POST /v1/completions
- Updated
- Reading time
- 6 min
- Level
- beginner
The native Revoye Cloud endpoint. Everything the API can do is expressible here — pick a model or let Revoye Cloud choose, wait for the answer or queue it, set your own timeouts, continue an earlier conversation, and attach metadata that comes back untouched.
curl $REVOYE_BASE_URL/v1/completions \
-H "Authorization: Bearer $REVOYE_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 4f1c9d2e-8b3a-4c1d-9e2f-7a6b5c4d3e2f" \
-d '{
"prompt": "Summarise the CAP theorem in three sentences.",
"provider": "chatgpt",
"wait": true
}'Requires the completions:write scope.
Request
| Field | Type | Default | Notes |
|---|---|---|---|
prompt | string, 1–100 000 chars | required | The text sent to the model |
provider | chatgpt | claude | gemini | deepseek | qwen | perplexity | null | null | null means any available model |
wait | boolean | true | Hold the HTTP request open until the job finishes |
timeout_ms | integer, 5 000–600 000 | 180 000 | Per attempt |
deadline_ms | integer, 10 000–3 600 000 | 900 000 | For the whole job, across all attempts |
priority | integer, −10…10 | 0 | Higher runs first within your own queue |
mode | new | continue | new | continue requires conversation_ref |
conversation_ref | string, ≤ 2 048 chars | null | Opaque; returned by a previous completion |
callback_url | https:// URL, ≤ 2 048 chars | null | Webhook when the job finishes |
metadata | object, ≤ 4 096 bytes as JSON | {} | Echoed back verbatim. Revoye Cloud never reads it |
agent_id | string | null | Advanced: restrict the job to the worker with this id, as reported in an earlier job's agent_id. Usually leave unset |
Headers. Idempotency-Key (at most 255 characters) is optional and strongly recommended — see
retrying safely.
Choosing a model, or not
Omitting provider is usually right: the job runs on whichever model has capacity first, which
gives you the shortest wait. Name a provider when the answer's quality depends on which model
produces it.
Naming one is a constraint, not a preference. Revoye Cloud will not silently answer a chatgpt
request from DeepSeek — receiving a different model than you asked for would be a correctness bug,
not a helpful fallback. The job waits for the model you named.
Continuing a conversation
Every completed job returns a conversation_ref. Send it back with mode: "continue" — and the same
provider — to ask a follow-up in the same conversation, so the model sees the earlier turns.
Treat the value as opaque: store it, do not parse or construct it.
metadata is yours
Whatever you put in comes back out, unchanged, on the job and on the webhook. Use it to correlate a job with your own records — a tenant id, a trace id, a row id — so that a webhook arriving an hour later can be routed without a database round trip.
The job object
Every endpoint on this page returns the same object.
{
"id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9",
"status": "succeeded",
"response": "The CAP theorem states that…",
"provider": "chatgpt",
"agent_id": "agt_01JAY7…",
"agent_name": null,
"conversation_ref": "<opaque string>",
"attempts": 1,
"queue_ms": 240,
"run_ms": 18432,
"created_at": "2026-09-18T09:14:02.000Z",
"finished_at": "2026-09-18T09:14:21.000Z",
"content_pruned_at": null,
"metadata": { "trace": "abc" }
}| Field | Means |
|---|---|
status | queued, dispatched (running), succeeded, failed, cancelled or expired |
response | The answer text. null until the job succeeds |
provider | The provider you asked for, or null if you let Revoye Cloud choose |
agent_id, agent_name | The worker that ran the job, for diagnostics. agent_name is currently always null |
attempts | How many times the job was started. > 1 means an earlier attempt failed or timed out |
queue_ms | How long it waited before its first attempt |
run_ms | Time spent actually generating the answer |
conversation_ref | Pass it back with mode: "continue" to continue the conversation |
content_pruned_at | When the prompt and response were removed under your account's retention setting. null while the text is still stored |
queue_position | Only on GET of a queued job: how many of your jobs are queued right now |
Status codes
| Request | Result |
|---|---|
wait: true, job succeeded | 200 with the job |
wait: false | 202 with the job, usually queued |
Repeated Idempotency-Key with wait: false | 200 with the original job |
wait: true, job did not succeed | An error: 502, 499 or 504, with details.job_id |
Blocking semantics
wait: true is a long poll, not synchronous execution. The job is recorded the moment it is
accepted, so:
- If your connection drops, the job keeps running.
- The result is at
GET /v1/completions/{id}, or delivered to yourcallback_url. - Nothing is lost by a client timeout.
The connection is held for at most min(deadline_ms, 600000) milliseconds — ten minutes with the
defaults. If that runs out first you get 504 JOB_TIMEOUT with details.job_id and
details.status, and the job continues.
Set your HTTP client's timeout above the hold time — 660 seconds covers the default — or do not use
wait: true. The single commonest integration bug is a 30-second default abandoning healthy work, and it looks exactly like the API failing, which it is not.
GET /v1/completions/{id}
Returns the job object with its current status. Poll it for a job you submitted with wait: false;
every few seconds is plenty.
curl "$REVOYE_BASE_URL/v1/completions/job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9" \
-H "Authorization: Bearer $REVOYE_API_KEY"A job that does not exist and a job that belongs to another account both return 404 NOT_FOUND.
Requires completions:read.
DELETE /v1/completions/{id}
Cancels a job and returns 200 with the job object. Requires completions:write.
- A
queuedjob is cancelled immediately and never runs. - A
dispatchedjob is markedcancelledstraight away. Generation that is already under way may finish, but its answer is discarded and never returned. - Cancelling a job that has already finished is a no-op, not an error: you get the job back as it is.
Limits
| Prompt | 100 000 characters |
| Request body | 1 MiB |
| Queued jobs per account | 1 000 |
metadata | 4 096 bytes, serialised as JSON |
Errors
| Code | HTTP | Retry? | Cause |
|---|---|---|---|
INVALID_REQUEST | 400 | No | details.field names the offender |
UNAUTHORIZED | 401 | No | Bad or revoked key |
FORBIDDEN | 403 | No | Key lacks the scope, or 1 000 of your jobs are already queued |
NOT_FOUND | 404 | No | No such job, or not yours |
CONFLICT | 409 | No | Idempotency key reused with a different request |
PAYLOAD_TOO_LARGE | 413 | No | Body over 1 MiB |
RATE_LIMITED | 429 | After Retry-After | Too many requests for this key |
NO_AGENT_AVAILABLE | 503 | Yes | No capacity for the model right now (only with wait: false and no callback_url) |
JOB_TIMEOUT | 504 | Yes | The hold ran out and the job continues — or the job's deadline passed |
JOB_FAILED | 502 | Maybe | Every attempt failed. The code may be more specific, such as PROMPT_REJECTED |
JOB_CANCELLED | 499 | No | The job was cancelled while you waited |
Full handling guidance: Errors.