An OpenAI API 429 response means the request was refused because a rate, token, quota, concurrency, or capacity constraint was reached. The status code alone is not enough to choose a fix. Read the error body, response headers, model, account, request size, and retry guidance together.

First response

Do not immediately send the same request again. Capture the request ID and error body, check balance and quota, inspect Retry-After, and determine whether the failed operation is safe to repeat.

Diagnose the 429

SignalLikely causeFirst action
Many requests in a short windowRequests-per-minute or concurrency limit.Queue work and lower parallelism.
Large prompts or outputsTokens-per-minute limit.Trim context and cap output.
Quota or billing languageAccount quota or balance exhausted.Check billing and account status; retries will not add quota.
One model fails while others workModel-specific limit or upstream capacity.Back off or use an approved fallback.
All models fail through a gatewayGateway group, account, or global quota.Inspect gateway usage and account limits.

Inspect the body and headers

const response = await fetch("https://api.sublyx.org/v1/chat/completions", options);

if (response.status === 429) {
  const body = await response.text();
  console.error({
    status: response.status,
    retryAfter: response.headers.get("retry-after"),
    requestId: response.headers.get("x-request-id"),
    body,
  });
}

Error field names can differ between the native provider and a compatible gateway. Treat the message as diagnostic input, not a stable programmatic contract. Use the HTTP status and your own retry policy as the primary control signal.

Safe exponential backoff

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

async function with429Retry(run, maxRetries = 4) {
  for (let attempt = 0; ; attempt += 1) {
    const response = await run();
    if (response.status !== 429 || attempt >= maxRetries) return response;

    const retryAfterSeconds = Number(response.headers.get("retry-after"));
    const retryAfter = Number.isFinite(retryAfterSeconds)
      ? retryAfterSeconds * 1000
      : Math.min(1000 * 2 ** attempt, 16000);
    const jitter = Math.random() * 500;
    await sleep(retryAfter + jitter);
  }
}

Only retry a request that is safe to repeat. A normal text completion may be repeated with extra cost and a different result. An asynchronous image or video submission may create a second task. Prefer an application idempotency key or a deduplication record for side-effecting work.

Python retry pattern

import random
import time
import requests

def post_with_backoff(url, *, headers, json, max_retries=4):
    for attempt in range(max_retries + 1):
        response = requests.post(url, headers=headers, json=json, timeout=(10, 120))
        if response.status_code != 429 or attempt == max_retries:
            return response

        retry_after = response.headers.get("retry-after")
        delay = float(retry_after) if retry_after else min(2 ** attempt, 16)
        time.sleep(delay + random.random() / 2)

    raise RuntimeError("unreachable")

RPM vs TPM

Requests per minute limits count calls. Tokens per minute limits the amount of input and output processed over time. A workload can stay below the request count while exceeding token throughput because every request contains a long conversation or requests a large completion.

  • Reduce repeated system instructions and old conversation turns.
  • Set a useful maximum output instead of an unlimited one.
  • Use a queue with separate concurrency budgets per model or tenant.
  • Cache deterministic results when the model, prompt, parameters, and version match.
  • Stagger scheduled jobs instead of releasing them in one burst.

OpenAI-compatible gateway checks

When using Sublyx, a 429 can be generated by the gateway group or passed through from an upstream model provider. Check the account balance, key group, model availability, recent request volume, concurrency, and request IDs. If only one model fails, inspect its route first; if every model fails, inspect account or global quota.

The existing general AI API 429 guide covers provider-neutral retry design. This page focuses on OpenAI and OpenAI-compatible request behavior. For routing, fallback, and cost controls, see the AI API gateway guide.

When not to retry

ResponseWhy a blind retry is wrong
400The request is malformed or uses an unsupported parameter.
401The credential must be fixed or rotated.
402Balance or payment state must be fixed.
Repeated 429 with quota wordingWaiting does not create more quota or balance.
Unknown side effectA replay may duplicate work or charge the user twice.

One-minute checklist

  1. Save the status, request ID, error body, model, and timestamp.
  2. Check Retry-After and rate-limit headers.
  3. Check account balance, group quota, RPM, TPM, and concurrency.
  4. Reduce load before retrying.
  5. Use exponential backoff with jitter and a maximum attempt count.
  6. Confirm the operation is safe to repeat.
  7. Switch models only when the fallback is tested and allowed.

Inspect your API limits

Use the Sublyx console to review keys, balance, model availability and usage before changing production retry behavior.

Read setup docsCompare modelsOpen console