An OpenAI API 429 response means the request was refused because a rate, token, quota, concurrency, or capacity constraint was reached. The status code alone is not enough to choose a fix. Read the error body, response headers, model, account, request size, and retry guidance together.
Do not immediately send the same request again. Capture the request ID and error body, check balance and quota, inspect Retry-After, and determine whether the failed operation is safe to repeat.
Diagnose the 429
| Signal | Likely cause | First action |
|---|---|---|
| Many requests in a short window | Requests-per-minute or concurrency limit. | Queue work and lower parallelism. |
| Large prompts or outputs | Tokens-per-minute limit. | Trim context and cap output. |
| Quota or billing language | Account quota or balance exhausted. | Check billing and account status; retries will not add quota. |
| One model fails while others work | Model-specific limit or upstream capacity. | Back off or use an approved fallback. |
| All models fail through a gateway | Gateway group, account, or global quota. | Inspect gateway usage and account limits. |
Inspect the body and headers
const response = await fetch("https://api.sublyx.org/v1/chat/completions", options);
if (response.status === 429) {
const body = await response.text();
console.error({
status: response.status,
retryAfter: response.headers.get("retry-after"),
requestId: response.headers.get("x-request-id"),
body,
});
}
Error field names can differ between the native provider and a compatible gateway. Treat the message as diagnostic input, not a stable programmatic contract. Use the HTTP status and your own retry policy as the primary control signal.
Safe exponential backoff
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
async function with429Retry(run, maxRetries = 4) {
for (let attempt = 0; ; attempt += 1) {
const response = await run();
if (response.status !== 429 || attempt >= maxRetries) return response;
const retryAfterSeconds = Number(response.headers.get("retry-after"));
const retryAfter = Number.isFinite(retryAfterSeconds)
? retryAfterSeconds * 1000
: Math.min(1000 * 2 ** attempt, 16000);
const jitter = Math.random() * 500;
await sleep(retryAfter + jitter);
}
}
Only retry a request that is safe to repeat. A normal text completion may be repeated with extra cost and a different result. An asynchronous image or video submission may create a second task. Prefer an application idempotency key or a deduplication record for side-effecting work.
Python retry pattern
import random
import time
import requests
def post_with_backoff(url, *, headers, json, max_retries=4):
for attempt in range(max_retries + 1):
response = requests.post(url, headers=headers, json=json, timeout=(10, 120))
if response.status_code != 429 or attempt == max_retries:
return response
retry_after = response.headers.get("retry-after")
delay = float(retry_after) if retry_after else min(2 ** attempt, 16)
time.sleep(delay + random.random() / 2)
raise RuntimeError("unreachable")
RPM vs TPM
Requests per minute limits count calls. Tokens per minute limits the amount of input and output processed over time. A workload can stay below the request count while exceeding token throughput because every request contains a long conversation or requests a large completion.
- Reduce repeated system instructions and old conversation turns.
- Set a useful maximum output instead of an unlimited one.
- Use a queue with separate concurrency budgets per model or tenant.
- Cache deterministic results when the model, prompt, parameters, and version match.
- Stagger scheduled jobs instead of releasing them in one burst.
OpenAI-compatible gateway checks
When using Sublyx, a 429 can be generated by the gateway group or passed through from an upstream model provider. Check the account balance, key group, model availability, recent request volume, concurrency, and request IDs. If only one model fails, inspect its route first; if every model fails, inspect account or global quota.
The existing general AI API 429 guide covers provider-neutral retry design. This page focuses on OpenAI and OpenAI-compatible request behavior. For routing, fallback, and cost controls, see the AI API gateway guide.
When not to retry
| Response | Why a blind retry is wrong |
|---|---|
400 | The request is malformed or uses an unsupported parameter. |
401 | The credential must be fixed or rotated. |
402 | Balance or payment state must be fixed. |
| Repeated 429 with quota wording | Waiting does not create more quota or balance. |
| Unknown side effect | A replay may duplicate work or charge the user twice. |
One-minute checklist
- Save the status, request ID, error body, model, and timestamp.
- Check
Retry-Afterand rate-limit headers. - Check account balance, group quota, RPM, TPM, and concurrency.
- Reduce load before retrying.
- Use exponential backoff with jitter and a maximum attempt count.
- Confirm the operation is safe to repeat.
- Switch models only when the fallback is tested and allowed.
Inspect your API limits
Use the Sublyx console to review keys, balance, model availability and usage before changing production retry behavior.
Read setup docsCompare modelsOpen console
Sublyx Field Notes