Rate limits and capacity
- Updated
- Reading time
- 3 min
- Level
- intermediate
Revoye Cloud has two kinds of limit and they behave differently: request limits refuse traffic that arrives too fast, while model capacity and hourly limits make work wait. Confusing them is the commonest source of "why is my integration slow" questions.
Request limits, per API key
Requests per minute, per API key. Exceeding them returns 429 RATE_LIMITED with a Retry-After
header in seconds. Honour the header.
| Limit | Covers | |
|---|---|---|
| Writes | 600 / minute | POST /v1/completions, POST /v1/chat/completions, DELETE /v1/completions/{id} |
| Reads | 1 200 / minute | GET /v1/completions/{id}, GET /v1/status, GET /v1/models |
Per key rather than per account, so one integration stuck in a retry loop cannot take the production key beside it down with it.
This limit is about request rate, not work volume. It is hard to hit with real prompts — each
takes tens of seconds to answer — and easy to hit by polling GET /v1/status or
GET /v1/completions/{id} in a tight loop. If you are seeing 429, look at your polling before you
look at your prompts.
Model capacity and hourly limits
Each model has a finite number of workers, and each worker handles one job at a time.
Some models also have an hourly limit on how many jobs they start. Both are set by Revoye Cloud and
shared by everyone using the service; GET /v1/status reports them per model as agents_idle,
rate_limit_per_hour and used_this_hour.
When a model has no free capacity, or its hourly limit is used up, your job does not fail — it
waits in the queue and starts when capacity frees up or the hour rolls over. Only a request with
wait: false and no callback_url is refused instead, with 503 NO_AGENT_AVAILABLE, because there
would be no way to tell you when it ran.
Your queue
At most 1 000 of your jobs can be queued at once. A request over that returns 403 FORBIDDEN
with details.limit. The queue is per account, across all your keys, and GET /v1/status reports
its depth and the age of its oldest job.
Within your own queue, a job with a higher priority (−10 to 10, default 0) starts first.
Planning throughput
A response typically takes 20 to 90 seconds, so throughput comes from running many jobs at once, not from any single one being quicker:
- Submit in parallel, collect asynchronously. Send jobs with
wait: falseand acallback_url, or hold severalwait: truerequests open at once. Do not send one prompt, wait, then send the next. - Let Revoye Cloud choose the model when you can. Omitting
providerlets a job start on whichever model is free first, instead of waiting for one in particular. - Watch
queue.oldest_queued_at. If it keeps getting older, you are submitting faster than the models you asked for can serve. Spread the work, or accept the delay.
Diagnosing a slow integration
| Symptom | Cause | Fix |
|---|---|---|
queue.depth climbing, agents_idle at zero for your model | That model is saturated | Omit provider, or accept the wait and use webhooks |
used_this_hour at rate_limit_per_hour | The model's hourly limit is used up | Jobs start when the hour rolls over; omit provider to use another model |
Frequent 429 RATE_LIMITED | You are polling too often | Poll every few seconds at most; use webhooks |
attempts > 1 on many jobs | Attempts are timing out | Raise timeout_ms for long prompts |
504 JOB_TIMEOUT with the job still running | Your wait: true hold ran out | Fetch the job later by id; nothing was lost |
Other limits
| Prompt | 100 000 characters |
messages on /v1/chat/completions | 256 messages, each ≤ 100 000 characters, flattening to ≤ 100 000 characters |
| Request body | 1 MiB |
| Queued jobs per account | 1 000 |
Per-attempt timeout, timeout_ms | 5 000–600 000 ms (default 180 000) |
Whole-job deadline, deadline_ms | 10 000–3 600 000 ms (default 900 000) |
wait: true hold | min(deadline_ms, 600 000) ms |
metadata | 4 096 bytes, serialised as JSON |
Idempotency-Key | 255 characters |