Skip to content

Best practices

Updated
Reading time
3 min
Level
intermediate

Design for a service where every request is a job and a typical answer takes 20 to 90 seconds. Every recommendation here follows from that one fact. An integration that respects it runs unattended for months; one that fights it feels unreliable.

1. Match the workload to the tool

Good fitPoor fit
Batch and scheduled workA chat UI with someone watching a cursor
Background enrichment, summarisation, classificationSub-second interactive features
Content pipelines with a deadline in minutes or hoursAnything with a deadline in seconds
Comparing answers from several models to the same promptStreaming output token by token

If a person is waiting on every keystroke, this is the wrong shape of API. That is not modesty — it is the difference between a happy integration and a support ticket.

2. Prefer asynchronous submission

wait: true is convenient for a script. For anything running unattended, submit with wait: false and a callback_url. It removes your dependency on a connection surviving a minute, removes the client-timeout bug class entirely, and turns "the model is busy" from an error into a delay. Receive results with a webhook.

3. Always send an idempotency key

Derive it from your own work item's identity, not from a fresh UUID inside the retry loop. Retrying safely.

4. Let Revoye Cloud choose the model when you can

Omit provider (or use revoye/auto) and each job runs on whichever model can take it first:

  • Throughput. A job does not wait for one busy model when another is free.
  • Resilience. If one model is saturated or at its hourly limit, work keeps flowing.

Pin a provider only when the answer's quality genuinely depends on which model produces it — and then expect to wait for that model's capacity.

5. Set timeouts deliberately

Prompt shapetimeout_ms
Short question, short answerThe default (180 000) is generous
Long document, long answer300 000–600 000
Research-style prompts that take the model a long timeToward the ceiling of 600 000

Set deadline_ms to the point at which the answer stops being useful to you. That is what stops a retry chain outliving its own purpose. And set your HTTP client's timeout above the wait: true hold — min(deadline_ms, 600000) — or do not hold the connection at all.

6. Write prompts for a chat model

Every provider answers as a conversational assistant. Two consequences:

  • Ask for the output format explicitly. "Reply with JSON only, no explanation." Chat models are chatty by default, and the API returns the whole answer as text.
  • Parse defensively. Strip code fences, tolerate a preamble, and validate before you trust. Treating the response as guaranteed-shaped JSON will break.

7. Instrument the four fields the job gives you

attempts, queue_ms, run_ms, provider. Log all four on every job and you can answer, without guessing:

  • Are jobs waiting (queue_ms high) or running slow (run_ms high)?
  • Is one model retrying more than the others (attempts by provider)?
  • Is your wait time changing over a week?

8. Monitor from /v1/status, on a schedule

Once a minute is plenty. Alert on your own queue.oldest_queued_at becoming older than your deadline, and watch agents_idle and used_this_hour for the models you depend on. Do not poll it per request — that is how you hit the read rate limit.

9. Store what you cannot recover

  • The job id. Record it as soon as you have it. With the id, any result can be fetched later; without it, a dropped connection means a lost answer.
  • conversation_ref, if you will ask follow-up questions. It is the only way to continue that conversation.
  • The response itself, if you need it long term. Prompt and response text are removed from Revoye Cloud after your account's retention period; content_pruned_at tells you when.

10. Keys: narrow, separate, rotatable

One key per deployment, scoped to what that deployment does, in a secret manager. Write-only for a submitter, read-only for a monitor. Authentication.