Rate limits
What CognitiveX actually throttles, what it does not, and how to handle 429 responses.
CognitiveX does not apply a global per-key request-rate limit. There is no
"N requests per second" ceiling on /api/recall, /api/remember, /api/talk,
or the other product endpoints, and the API does not emit X-RateLimit-*
headers anywhere.
What does limit you is one of three things, each of which surfaces as HTTP 429 Too Many Requests with a distinct body:
- A small fixed rate limit on authentication endpoints (sign-up, sign-in, password reset), to slow credential-stuffing.
- Your monthly tier quota (for example recall credits per month), which is a usage cap, not a request-rate throttle.
- Your pay-as-you-go spend cap running out, on the
paygtier.
This page documents all three. For the quota numbers themselves and how to raise them, see Billing and usage.
There is no Idempotency-Key header and no Stripe-style idempotency on this
API. Do not assume a retried POST is deduplicated for you. The only
rate-limit-adjacent header anywhere is Retry-After, and it is only sent on
the /api/url extraction limiter (covered below), not on the auth limiter or
on quota errors.
Auth endpoint rate limit
The auth endpoints are throttled to 5 requests per 60 seconds, per client IP (not per user, not per key). It is a fixed in-memory limiter, so it resets after the 60-second window passes.
It applies only to these endpoints:
| Endpoint | Method |
|---|---|
/api/auth/signup | POST |
/api/auth/signin | POST |
/api/auth/forgot-password | POST |
/api/auth/reset-password | POST |
/api/auth/resend-verification | POST |
When you exceed it, you get:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
{ "detail": "Too many requests. Please try again later." }There is no Retry-After header on this response. Wait out the 60-second
window and retry. In practice you only hit this with retry loops or load tests
against sign-in; normal interactive auth never approaches 5 attempts/minute.
Tier quotas
Most metered usage is gated by your monthly tier quota rather than by request
rate. The tiers are amnesiac (free), awakened, conscious, and payg. When
a metered resource (for example recall) exceeds your tier's monthly allowance,
the request returns 429 with a structured body:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
{
"detail": {
"error": "quota_exceeded",
"message": "You've reached your monthly limit of 100 recalls. Upgrade to Awakened tier for more.",
"quota": { "used": 100, "limit": 100, "remaining": 0 },
"current_tier": "amnesiac",
"recommended_tier": "awakened",
"upgrade_url": "/settings"
}
}Detect a quota error by checking that detail is an object with
error === "quota_exceeded" (as opposed to the auth limiter, whose detail is a
plain string). Read detail.quota.limit and detail.recommended_tier to decide
whether to surface an upgrade prompt or back off until the next monthly reset.
The quota numbers per tier (recall credits per month, memory cap, and so on) live
in Billing and usage. They are also served
live from GET /api/billing/tiers, which is the authoritative source if the
docs and the running system ever disagree.
Pay-as-you-go spend caps
On the payg tier there is no monthly message quota. Instead, usage is metered
against credits and bounded by an optional spend cap you set yourself. When a
request would exceed the cap, it is rejected the same way a quota would be (429
with a structured detail), so handle it with the same code path as a tier
quota error.
Set or read your cap through the billing surface:
| Endpoint | Method | Purpose |
|---|---|---|
/api/billing/payg-cap | GET | Read the current spend cap. |
/api/billing/payg-cap | PUT | Set or clear the spend cap. |
/api/billing/balance | GET | Read the current credit balance. |
See Billing and usage for the full request and response shapes. Note that credits are the user-facing billing unit; USD figures on these endpoints are billing detail.
The /api/url extraction limiter
The one endpoint that does carry a per-user request limiter with a Retry-After
header is /api/url (URL fetch and extract). It is the only place in the API
where you should read Retry-After to schedule a retry:
HTTP/1.1 429 Too Many Requests
Retry-After: 12
Content-Type: application/json
{ "error": "rate_limited", "retry_after": 12, "method": "extract" }Wait retry_after seconds (or read the Retry-After header, which carries the
same value) before retrying that method. This shape is specific to /api/url;
do not expect it elsewhere.
Handling 429 in client code
Because the three 429 causes have different bodies, branch on the shape of
detail (or error) rather than treating every 429 the same:
detailis a string ("Too many requests..."): the auth limiter. Wait ~60s, retry. NoRetry-After.detailis an object witherror: "quota_exceeded": a tier quota or PAYG cap. Retrying immediately will not help; the user must upgrade, raise their spend cap, or wait for the monthly reset. Surfacedetail.recommended_tier/detail.upgrade_url.- A top-level
error: "rate_limited"withretry_after(only from/api/url): honorretry_after/ theRetry-Afterheader, then retry.
curl -i -X POST https://api.cognitivx.io/api/recall \
-H "Authorization: Bearer icog_xxx" \
-H "Content-Type: application/json" \
-d '{"query":"what did we decide about billing","limit":5}'A 429 here means you hit your recall quota, not a request-rate throttle. Inspect
the JSON body to confirm and read recommended_tier.
import { configureApiClient, memory, ApiError } from '@cognitivx/sdk';
configureApiClient({ apiKey: 'icog_xxx' });
try {
const res = await memory.recall({ query: 'billing decisions', limit: 5 });
console.log(res.memories);
} catch (err) {
if (err instanceof ApiError && err.status === 429) {
const detail = (err.body as { detail?: unknown })?.detail;
if (detail && typeof detail === 'object' && 'error' in detail) {
// quota_exceeded: do not retry, prompt an upgrade
console.error('Quota reached:', detail);
} else {
// auth limiter: back off ~60s and retry
console.error('Rate limited, retry shortly');
}
} else {
throw err;
}
}The SDK does not silently retry 429s for you. Quota and spend-cap 429s are not retryable (the answer is upgrade, raise the cap, or wait for the monthly reset), so automatic retry would only burn requests. Branch on the body and decide explicitly.
Tracking usage
There is no console.cognitivx.io usage dashboard. Read your usage and balance
programmatically:
| Endpoint | Method | Purpose |
|---|---|---|
/api/billing/usage | GET | Current period usage against your quota. |
/api/billing/balance | GET | PAYG credit balance (balance_credits, balance_usd, low_balance, depleted). |
/api/billing/me | GET | Tier, subscription state, and a PAYG snapshot. |
See Billing and usage for the full surface, and Authentication for how to pass your key.