Operations

Rate limits

Per-key limits by plan (60, 180 or 600 requests a minute), concurrency and word throughput, and how to back off.

Limits apply per API key, so a runaway batch job on one key can't starve your production traffic on another. They follow your plan.

Limits by plan

PlanRequests / minConcurrent requestsWords / minAsync jobs in flightVoices
Basic API605150k503
Pro API18020600k50015
Ultra API600502M5,000Unlimited
Test keys (sk_test_)60550k20Unlimited

Deep strength and calibrated Detect need Pro API or Ultra API (403 permission_denied otherwise). Queued jobs run in plan order: Ultra API first.

Rate limit headers

Every response tells you where you stand, so you can slow down before you hit the wall:

HeaderMeaning
X-RateLimit-LimitRequests allowed in the current one-minute window.
X-RateLimit-RemainingRequests left in the window.
X-RateLimit-ResetUnix time when the window resets.
X-WordLimit-RemainingWords left in the current word-throughput window.
Retry-AfterSeconds to wait. Only on 429.

Handling 429

Respect Retry-After when present; otherwise back off exponentially with jitter. Rejected requests are never billed.

async function withRetry(call: () => Promise<Response>, max = 5) {
  for (let attempt = 0; ; attempt++) {
    const res = await call();
    if (res.status !== 429 && res.status < 500) return res;
    if (attempt >= max) return res;
    const retryAfter = Number(res.headers.get("retry-after"));
    const wait = retryAfter > 0 ? retryAfter * 1000 : Math.min(2 ** attempt * 500, 20000) * (0.5 + Math.random());
    await new Promise((r) => setTimeout(r, wait));
  }
}

Tip:

For large batches, use async jobs: they queue on our side instead of competing for your per-minute request budget.