Best practices
- Updated
- Reading time
- 3 min
- Level
- intermediate
On this page
- 1. Match the workload to the tool
- 2. Prefer asynchronous submission
- 3. Always send an idempotency key
- 4. Let Revoye Cloud choose the model when you can
- 5. Set timeouts deliberately
- 6. Write prompts for a chat model
- 7. Instrument the four fields the job gives you
- 8. Monitor from /v1/status, on a schedule
- 9. Store what you cannot recover
- 10. Keys: narrow, separate, rotatable
Design for a service where every request is a job and a typical answer takes 20 to 90 seconds. Every recommendation here follows from that one fact. An integration that respects it runs unattended for months; one that fights it feels unreliable.
1. Match the workload to the tool
| Good fit | Poor fit |
|---|---|
| Batch and scheduled work | A chat UI with someone watching a cursor |
| Background enrichment, summarisation, classification | Sub-second interactive features |
| Content pipelines with a deadline in minutes or hours | Anything with a deadline in seconds |
| Comparing answers from several models to the same prompt | Streaming output token by token |
If a person is waiting on every keystroke, this is the wrong shape of API. That is not modesty — it is the difference between a happy integration and a support ticket.
2. Prefer asynchronous submission
wait: true is convenient for a script. For anything running unattended, submit with wait: false
and a callback_url. It removes your dependency on a connection surviving a minute, removes the
client-timeout bug class entirely, and turns "the model is busy" from an error into a delay.
Receive results with a webhook.
3. Always send an idempotency key
Derive it from your own work item's identity, not from a fresh UUID inside the retry loop. Retrying safely.
4. Let Revoye Cloud choose the model when you can
Omit provider (or use revoye/auto) and each job runs on whichever model can take it first:
- Throughput. A job does not wait for one busy model when another is free.
- Resilience. If one model is saturated or at its hourly limit, work keeps flowing.
Pin a provider only when the answer's quality genuinely depends on which model produces it — and then expect to wait for that model's capacity.
5. Set timeouts deliberately
| Prompt shape | timeout_ms |
|---|---|
| Short question, short answer | The default (180 000) is generous |
| Long document, long answer | 300 000–600 000 |
| Research-style prompts that take the model a long time | Toward the ceiling of 600 000 |
Set deadline_ms to the point at which the answer stops being useful to you. That is what stops a
retry chain outliving its own purpose. And set your HTTP client's timeout above the wait: true hold
— min(deadline_ms, 600000) — or do not hold the connection at all.
6. Write prompts for a chat model
Every provider answers as a conversational assistant. Two consequences:
- Ask for the output format explicitly. "Reply with JSON only, no explanation." Chat models are chatty by default, and the API returns the whole answer as text.
- Parse defensively. Strip code fences, tolerate a preamble, and validate before you trust. Treating the response as guaranteed-shaped JSON will break.
7. Instrument the four fields the job gives you
attempts, queue_ms, run_ms, provider. Log all four on every job and you can answer, without
guessing:
- Are jobs waiting (
queue_mshigh) or running slow (run_mshigh)? - Is one model retrying more than the others (
attemptsbyprovider)? - Is your wait time changing over a week?
8. Monitor from /v1/status, on a schedule
Once a minute is plenty. Alert on your own queue.oldest_queued_at becoming older than your deadline,
and watch agents_idle and used_this_hour for the models you depend on. Do not poll it per request
— that is how you hit the read rate limit.
9. Store what you cannot recover
- The job id. Record it as soon as you have it. With the id, any result can be fetched later; without it, a dropped connection means a lost answer.
conversation_ref, if you will ask follow-up questions. It is the only way to continue that conversation.- The response itself, if you need it long term. Prompt and response text are removed from
Revoye Cloud after your account's retention period;
content_pruned_attells you when.
10. Keys: narrow, separate, rotatable
One key per deployment, scoped to what that deployment does, in a secret manager. Write-only for a submitter, read-only for a monitor. Authentication.