Skip to main content
The Big Brain Ape API enforces rate limits to keep the platform reliable and fair for all users. Limits are applied per API key on a rolling one-minute window, and a separate daily quota caps total usage over a 24-hour period. If your integration exceeds either limit, the API returns a 429 Too Many Requests response until the current window resets. Understanding these limits — and designing your code to respect them — keeps your integration running smoothly without unnecessary interruptions.

Rate limits by plan

Limits apply to the API key making the request, not to your account as a whole. If you generate multiple keys under the same plan, each key receives its own independent limit. To increase your limits, upgrade your plan from the Settings → Billing page in the dashboard.

Rate limit headers

Every API response includes headers that tell you exactly where you stand in the current window. There are two sets: one for the per-minute window and one for the daily quota. Per-minute window headers: Daily quota headers: Read these headers in your integration to implement proactive throttling — for example, slow down your request cadence when X-RateLimit-Remaining drops below a threshold rather than waiting for a 429.

Handling 429 errors

When you exceed your rate limit, the API responds with:
The response body follows the standard error format:
The response also includes a Retry-After header containing the number of seconds you should wait before sending another request. Always honour this value rather than retrying immediately — hammering the API while rate-limited does not help and may result in longer back-off periods.

Implementing retry with backoff

Use exponential backoff to handle 429 responses gracefully. The following example retries the request up to three times, honouring the Retry-After header when present and falling back to an exponential delay otherwise:
Apply the same pattern in any language: check the status code, read Retry-After, sleep for that duration, then retry. Cap the number of retries and surface an error to the caller if all attempts are exhausted.
Reduce the number of requests your integration makes by batching where possible — for example, fetch data for multiple tokens in a single request instead of making one call per token. For frequently read values like token prices, cache responses locally for a short window (prices are valid for approximately 5 seconds) to avoid redundant calls that consume your quota without returning fresher data.