CognitiveX Docs

Rate limits

What CognitiveX actually throttles, what it does not, and how to handle 429 responses.

CognitiveX does not apply a global per-key request-rate limit. There is no "N requests per second" ceiling on /api/recall, /api/remember, /api/talk, or the other product endpoints, and the API does not emit X-RateLimit-* headers anywhere.

What does limit you is one of three things, each of which surfaces as HTTP 429 Too Many Requests with a distinct body:

  1. A small fixed rate limit on authentication endpoints (sign-up, sign-in, password reset), to slow credential-stuffing.
  2. Your monthly tier quota (for example recall credits per month), which is a usage cap, not a request-rate throttle.
  3. Your pay-as-you-go spend cap running out, on the payg tier.

This page documents all three. For the quota numbers themselves and how to raise them, see Billing and usage.

There is no Idempotency-Key header and no Stripe-style idempotency on this API. Do not assume a retried POST is deduplicated for you. The only rate-limit-adjacent header anywhere is Retry-After, and it is only sent on the /api/url extraction limiter (covered below), not on the auth limiter or on quota errors.

Auth endpoint rate limit

The auth endpoints are throttled to 5 requests per 60 seconds, per client IP (not per user, not per key). It is a fixed in-memory limiter, so it resets after the 60-second window passes.

It applies only to these endpoints:

EndpointMethod
/api/auth/signupPOST
/api/auth/signinPOST
/api/auth/forgot-passwordPOST
/api/auth/reset-passwordPOST
/api/auth/resend-verificationPOST

When you exceed it, you get:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{ "detail": "Too many requests. Please try again later." }

There is no Retry-After header on this response. Wait out the 60-second window and retry. In practice you only hit this with retry loops or load tests against sign-in; normal interactive auth never approaches 5 attempts/minute.

Tier quotas

Most metered usage is gated by your monthly tier quota rather than by request rate. The tiers are amnesiac (free), awakened, conscious, and payg. When a metered resource (for example recall) exceeds your tier's monthly allowance, the request returns 429 with a structured body:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{
  "detail": {
    "error": "quota_exceeded",
    "message": "You've reached your monthly limit of 100 recalls. Upgrade to Awakened tier for more.",
    "quota": { "used": 100, "limit": 100, "remaining": 0 },
    "current_tier": "amnesiac",
    "recommended_tier": "awakened",
    "upgrade_url": "/settings"
  }
}

Detect a quota error by checking that detail is an object with error === "quota_exceeded" (as opposed to the auth limiter, whose detail is a plain string). Read detail.quota.limit and detail.recommended_tier to decide whether to surface an upgrade prompt or back off until the next monthly reset.

The quota numbers per tier (recall credits per month, memory cap, and so on) live in Billing and usage. They are also served live from GET /api/billing/tiers, which is the authoritative source if the docs and the running system ever disagree.

Pay-as-you-go spend caps

On the payg tier there is no monthly message quota. Instead, usage is metered against credits and bounded by an optional spend cap you set yourself. When a request would exceed the cap, it is rejected the same way a quota would be (429 with a structured detail), so handle it with the same code path as a tier quota error.

Set or read your cap through the billing surface:

EndpointMethodPurpose
/api/billing/payg-capGETRead the current spend cap.
/api/billing/payg-capPUTSet or clear the spend cap.
/api/billing/balanceGETRead the current credit balance.

See Billing and usage for the full request and response shapes. Note that credits are the user-facing billing unit; USD figures on these endpoints are billing detail.

The /api/url extraction limiter

The one endpoint that does carry a per-user request limiter with a Retry-After header is /api/url (URL fetch and extract). It is the only place in the API where you should read Retry-After to schedule a retry:

HTTP/1.1 429 Too Many Requests
Retry-After: 12
Content-Type: application/json

{ "error": "rate_limited", "retry_after": 12, "method": "extract" }

Wait retry_after seconds (or read the Retry-After header, which carries the same value) before retrying that method. This shape is specific to /api/url; do not expect it elsewhere.

Handling 429 in client code

Because the three 429 causes have different bodies, branch on the shape of detail (or error) rather than treating every 429 the same:

  • detail is a string ("Too many requests..."): the auth limiter. Wait ~60s, retry. No Retry-After.
  • detail is an object with error: "quota_exceeded": a tier quota or PAYG cap. Retrying immediately will not help; the user must upgrade, raise their spend cap, or wait for the monthly reset. Surface detail.recommended_tier / detail.upgrade_url.
  • A top-level error: "rate_limited" with retry_after (only from /api/url): honor retry_after / the Retry-After header, then retry.
curl -i -X POST https://api.cognitivx.io/api/recall \
  -H "Authorization: Bearer icog_xxx" \
  -H "Content-Type: application/json" \
  -d '{"query":"what did we decide about billing","limit":5}'

A 429 here means you hit your recall quota, not a request-rate throttle. Inspect the JSON body to confirm and read recommended_tier.

import { configureApiClient, memory, ApiError } from '@cognitivx/sdk';

configureApiClient({ apiKey: 'icog_xxx' });

try {
  const res = await memory.recall({ query: 'billing decisions', limit: 5 });
  console.log(res.memories);
} catch (err) {
  if (err instanceof ApiError && err.status === 429) {
    const detail = (err.body as { detail?: unknown })?.detail;
    if (detail && typeof detail === 'object' && 'error' in detail) {
      // quota_exceeded: do not retry, prompt an upgrade
      console.error('Quota reached:', detail);
    } else {
      // auth limiter: back off ~60s and retry
      console.error('Rate limited, retry shortly');
    }
  } else {
    throw err;
  }
}

The SDK does not silently retry 429s for you. Quota and spend-cap 429s are not retryable (the answer is upgrade, raise the cap, or wait for the monthly reset), so automatic retry would only burn requests. Branch on the body and decide explicitly.

Tracking usage

There is no console.cognitivx.io usage dashboard. Read your usage and balance programmatically:

EndpointMethodPurpose
/api/billing/usageGETCurrent period usage against your quota.
/api/billing/balanceGETPAYG credit balance (balance_credits, balance_usd, low_balance, depleted).
/api/billing/meGETTier, subscription state, and a PAYG snapshot.

See Billing and usage for the full surface, and Authentication for how to pass your key.