Operations
Rate limits
Per-key limits by plan (60, 180 or 600 requests a minute), concurrency and word throughput, and how to back off.
Limits apply per API key, so a runaway batch job on one key can't starve your production traffic on another. They follow your plan.
Limits by plan
| Plan | Requests / min | Concurrent requests | Words / min | Async jobs in flight | Voices |
|---|---|---|---|---|---|
| Basic API | 60 | 5 | 150k | 50 | 3 |
| Pro API | 180 | 20 | 600k | 500 | 15 |
| Ultra API | 600 | 50 | 2M | 5,000 | Unlimited |
Test keys (sk_test_) | 60 | 5 | 50k | 20 | Unlimited |
Deep strength and calibrated Detect need Pro API or Ultra API (403 permission_denied otherwise). Queued jobs run in plan order: Ultra API first.
Rate limit headers
Every response tells you where you stand, so you can slow down before you hit the wall:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Requests allowed in the current one-minute window. |
X-RateLimit-Remaining | Requests left in the window. |
X-RateLimit-Reset | Unix time when the window resets. |
X-WordLimit-Remaining | Words left in the current word-throughput window. |
Retry-After | Seconds to wait. Only on 429. |
Handling 429
Respect Retry-After when present; otherwise back off exponentially with jitter. Rejected requests are never billed.
async function withRetry(call: () => Promise<Response>, max = 5) {
for (let attempt = 0; ; attempt++) {
const res = await call();
if (res.status !== 429 && res.status < 500) return res;
if (attempt >= max) return res;
const retryAfter = Number(res.headers.get("retry-after"));
const wait = retryAfter > 0 ? retryAfter * 1000 : Math.min(2 ** attempt * 500, 20000) * (0.5 + Math.random());
await new Promise((r) => setTimeout(r, wait));
}
}Tip:
For large batches, use async jobs: they queue on our side instead of competing for your per-minute request budget.