Skip to content

Rate limits and capacity

Updated
Reading time
3 min
Level
intermediate

Revoye Cloud has two kinds of limit and they behave differently: request limits refuse traffic that arrives too fast, while model capacity and hourly limits make work wait. Confusing them is the commonest source of "why is my integration slow" questions.

Request limits, per API key

Requests per minute, per API key. Exceeding them returns 429 RATE_LIMITED with a Retry-After header in seconds. Honour the header.

LimitCovers
Writes600 / minutePOST /v1/completions, POST /v1/chat/completions, DELETE /v1/completions/{id}
Reads1 200 / minuteGET /v1/completions/{id}, GET /v1/status, GET /v1/models

Per key rather than per account, so one integration stuck in a retry loop cannot take the production key beside it down with it.

This limit is about request rate, not work volume. It is hard to hit with real prompts — each takes tens of seconds to answer — and easy to hit by polling GET /v1/status or GET /v1/completions/{id} in a tight loop. If you are seeing 429, look at your polling before you look at your prompts.

Model capacity and hourly limits

Each model has a finite number of workers, and each worker handles one job at a time. Some models also have an hourly limit on how many jobs they start. Both are set by Revoye Cloud and shared by everyone using the service; GET /v1/status reports them per model as agents_idle, rate_limit_per_hour and used_this_hour.

When a model has no free capacity, or its hourly limit is used up, your job does not fail — it waits in the queue and starts when capacity frees up or the hour rolls over. Only a request with wait: false and no callback_url is refused instead, with 503 NO_AGENT_AVAILABLE, because there would be no way to tell you when it ran.

Your queue

At most 1 000 of your jobs can be queued at once. A request over that returns 403 FORBIDDEN with details.limit. The queue is per account, across all your keys, and GET /v1/status reports its depth and the age of its oldest job.

Within your own queue, a job with a higher priority (−10 to 10, default 0) starts first.

Planning throughput

A response typically takes 20 to 90 seconds, so throughput comes from running many jobs at once, not from any single one being quicker:

  • Submit in parallel, collect asynchronously. Send jobs with wait: false and a callback_url, or hold several wait: true requests open at once. Do not send one prompt, wait, then send the next.
  • Let Revoye Cloud choose the model when you can. Omitting provider lets a job start on whichever model is free first, instead of waiting for one in particular.
  • Watch queue.oldest_queued_at. If it keeps getting older, you are submitting faster than the models you asked for can serve. Spread the work, or accept the delay.

Diagnosing a slow integration

SymptomCauseFix
queue.depth climbing, agents_idle at zero for your modelThat model is saturatedOmit provider, or accept the wait and use webhooks
used_this_hour at rate_limit_per_hourThe model's hourly limit is used upJobs start when the hour rolls over; omit provider to use another model
Frequent 429 RATE_LIMITEDYou are polling too oftenPoll every few seconds at most; use webhooks
attempts > 1 on many jobsAttempts are timing outRaise timeout_ms for long prompts
504 JOB_TIMEOUT with the job still runningYour wait: true hold ran outFetch the job later by id; nothing was lost

Other limits

Prompt100 000 characters
messages on /v1/chat/completions256 messages, each ≤ 100 000 characters, flattening to ≤ 100 000 characters
Request body1 MiB
Queued jobs per account1 000
Per-attempt timeout, timeout_ms5 000–600 000 ms (default 180 000)
Whole-job deadline, deadline_ms10 000–3 600 000 ms (default 900 000)
wait: true holdmin(deadline_ms, 600 000) ms
metadata4 096 bytes, serialised as JSON
Idempotency-Key255 characters