> ## Documentation Index
> Fetch the complete documentation index at: https://bigbrainape.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Big Brain Ape API Rate Limits, Quotas, and Backoff Strategies

> Understand Big Brain Ape API rate limits by plan, how to read rate limit headers, handle 429 errors, and implement backoff strategies.

The Big Brain Ape API enforces rate limits to keep the platform reliable and fair for all users. Limits are applied per API key on a rolling one-minute window, and a separate daily quota caps total usage over a 24-hour period. If your integration exceeds either limit, the API returns a `429 Too Many Requests` response until the current window resets. Understanding these limits — and designing your code to respect them — keeps your integration running smoothly without unnecessary interruptions.

## Rate limits by plan

| Plan | Requests per minute | Requests per day |
| - | - | - |
| Free | 60 | 5,000 |
| Pro | 300 | 50,000 |

Limits apply to the API key making the request, not to your account as a whole. If you generate multiple keys under the same plan, each key receives its own independent limit. To increase your limits, upgrade your plan from the **Settings → Billing** page in the dashboard.

## Rate limit headers

Every API response includes headers that tell you exactly where you stand in the current window. There are two sets: one for the per-minute window and one for the daily quota.

**Per-minute window headers:**

| Header | Description |
| - | - |
| `X-RateLimit-Limit` | Your total request allowance for the current one-minute window |
| `X-RateLimit-Remaining` | Number of requests you can still make before the window resets |
| `X-RateLimit-Reset` | Unix timestamp (UTC) at which the current window expires and your limit refreshes |

**Daily quota headers:**

| Header | Description |
| - | - |
| `X-RateLimit-Daily-Limit` | Your total request allowance for the current 24-hour period |
| `X-RateLimit-Daily-Remaining` | Number of requests remaining in the current 24-hour period |
| `X-RateLimit-Daily-Reset` | Unix timestamp (UTC) at which the daily quota resets |

Read these headers in your integration to implement proactive throttling — for example, slow down your request cadence when `X-RateLimit-Remaining` drops below a threshold rather than waiting for a `429`.

## Handling 429 errors

When you exceed your rate limit, the API responds with:

```
HTTP 429 Too Many Requests
```

The response body follows the standard error format:

```json theme={null}
{
  "success": false,
  "error": {
    "code": "RATE_LIMIT_EXCEEDED",
    "message": "You have exceeded your request limit. Please wait before retrying."
  }
}
```

The response also includes a `Retry-After` header containing the number of seconds you should wait before sending another request. Always honour this value rather than retrying immediately — hammering the API while rate-limited does not help and may result in longer back-off periods.

## Implementing retry with backoff

Use exponential backoff to handle `429` responses gracefully. The following example retries the request up to three times, honouring the `Retry-After` header when present and falling back to an exponential delay otherwise:

```python theme={null}
import time
import requests

def api_request_with_backoff(url, headers, max_retries=3):
    for attempt in range(max_retries):
        response = requests.get(url, headers=headers)
        if response.status_code == 429:
            retry_after = int(response.headers.get("Retry-After", 2 ** attempt))
            time.sleep(retry_after)
            continue
        return response
    raise Exception("Max retries exceeded")
```

Apply the same pattern in any language: check the status code, read `Retry-After`, sleep for that duration, then retry. Cap the number of retries and surface an error to the caller if all attempts are exhausted.

<Tip>
  Reduce the number of requests your integration makes by batching where possible — for example, fetch data for multiple tokens in a single request instead of making one call per token. For frequently read values like token prices, cache responses locally for a short window (prices are valid for approximately 5 seconds) to avoid redundant calls that consume your quota without returning fresher data.
</Tip>
