Skip to content

POST /v1/completions

Updated
Reading time
6 min
Level
beginner

The native Revoye Cloud endpoint. Everything the API can do is expressible here — pick a model or let Revoye Cloud choose, wait for the answer or queue it, set your own timeouts, continue an earlier conversation, and attach metadata that comes back untouched.

curl $REVOYE_BASE_URL/v1/completions \
  -H "Authorization: Bearer $REVOYE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: 4f1c9d2e-8b3a-4c1d-9e2f-7a6b5c4d3e2f" \
  -d '{
    "prompt": "Summarise the CAP theorem in three sentences.",
    "provider": "chatgpt",
    "wait": true
  }'

Requires the completions:write scope.

Request

FieldTypeDefaultNotes
promptstring, 1–100 000 charsrequiredThe text sent to the model
providerchatgpt | claude | gemini | deepseek | qwen | perplexity | nullnullnull means any available model
waitbooleantrueHold the HTTP request open until the job finishes
timeout_msinteger, 5 000–600 000180 000Per attempt
deadline_msinteger, 10 000–3 600 000900 000For the whole job, across all attempts
priorityinteger, −10…100Higher runs first within your own queue
modenew | continuenewcontinue requires conversation_ref
conversation_refstring, ≤ 2 048 charsnullOpaque; returned by a previous completion
callback_urlhttps:// URL, ≤ 2 048 charsnullWebhook when the job finishes
metadataobject, ≤ 4 096 bytes as JSON{}Echoed back verbatim. Revoye Cloud never reads it
agent_idstringnullAdvanced: restrict the job to the worker with this id, as reported in an earlier job's agent_id. Usually leave unset

Headers. Idempotency-Key (at most 255 characters) is optional and strongly recommended — see retrying safely.

Choosing a model, or not

Omitting provider is usually right: the job runs on whichever model has capacity first, which gives you the shortest wait. Name a provider when the answer's quality depends on which model produces it.

Naming one is a constraint, not a preference. Revoye Cloud will not silently answer a chatgpt request from DeepSeek — receiving a different model than you asked for would be a correctness bug, not a helpful fallback. The job waits for the model you named.

Continuing a conversation

Every completed job returns a conversation_ref. Send it back with mode: "continue" — and the same provider — to ask a follow-up in the same conversation, so the model sees the earlier turns. Treat the value as opaque: store it, do not parse or construct it.

metadata is yours

Whatever you put in comes back out, unchanged, on the job and on the webhook. Use it to correlate a job with your own records — a tenant id, a trace id, a row id — so that a webhook arriving an hour later can be routed without a database round trip.

The job object

Every endpoint on this page returns the same object.

{
  "id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9",
  "status": "succeeded",
  "response": "The CAP theorem states that…",
  "provider": "chatgpt",
  "agent_id": "agt_01JAY7…",
  "agent_name": null,
  "conversation_ref": "<opaque string>",
  "attempts": 1,
  "queue_ms": 240,
  "run_ms": 18432,
  "created_at": "2026-09-18T09:14:02.000Z",
  "finished_at": "2026-09-18T09:14:21.000Z",
  "content_pruned_at": null,
  "metadata": { "trace": "abc" }
}
FieldMeans
statusqueued, dispatched (running), succeeded, failed, cancelled or expired
responseThe answer text. null until the job succeeds
providerThe provider you asked for, or null if you let Revoye Cloud choose
agent_id, agent_nameThe worker that ran the job, for diagnostics. agent_name is currently always null
attemptsHow many times the job was started. > 1 means an earlier attempt failed or timed out
queue_msHow long it waited before its first attempt
run_msTime spent actually generating the answer
conversation_refPass it back with mode: "continue" to continue the conversation
content_pruned_atWhen the prompt and response were removed under your account's retention setting. null while the text is still stored
queue_positionOnly on GET of a queued job: how many of your jobs are queued right now

Status codes

RequestResult
wait: true, job succeeded200 with the job
wait: false202 with the job, usually queued
Repeated Idempotency-Key with wait: false200 with the original job
wait: true, job did not succeedAn error: 502, 499 or 504, with details.job_id

Blocking semantics

wait: true is a long poll, not synchronous execution. The job is recorded the moment it is accepted, so:

  • If your connection drops, the job keeps running.
  • The result is at GET /v1/completions/{id}, or delivered to your callback_url.
  • Nothing is lost by a client timeout.

The connection is held for at most min(deadline_ms, 600000) milliseconds — ten minutes with the defaults. If that runs out first you get 504 JOB_TIMEOUT with details.job_id and details.status, and the job continues.

Set your HTTP client's timeout above the hold time — 660 seconds covers the default — or do not use wait: true. The single commonest integration bug is a 30-second default abandoning healthy work, and it looks exactly like the API failing, which it is not.

GET /v1/completions/{id}

Returns the job object with its current status. Poll it for a job you submitted with wait: false; every few seconds is plenty.

curl "$REVOYE_BASE_URL/v1/completions/job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9" \
  -H "Authorization: Bearer $REVOYE_API_KEY"

A job that does not exist and a job that belongs to another account both return 404 NOT_FOUND. Requires completions:read.

DELETE /v1/completions/{id}

Cancels a job and returns 200 with the job object. Requires completions:write.

  • A queued job is cancelled immediately and never runs.
  • A dispatched job is marked cancelled straight away. Generation that is already under way may finish, but its answer is discarded and never returned.
  • Cancelling a job that has already finished is a no-op, not an error: you get the job back as it is.

Limits

Prompt100 000 characters
Request body1 MiB
Queued jobs per account1 000
metadata4 096 bytes, serialised as JSON

Errors

CodeHTTPRetry?Cause
INVALID_REQUEST400Nodetails.field names the offender
UNAUTHORIZED401NoBad or revoked key
FORBIDDEN403NoKey lacks the scope, or 1 000 of your jobs are already queued
NOT_FOUND404NoNo such job, or not yours
CONFLICT409NoIdempotency key reused with a different request
PAYLOAD_TOO_LARGE413NoBody over 1 MiB
RATE_LIMITED429After Retry-AfterToo many requests for this key
NO_AGENT_AVAILABLE503YesNo capacity for the model right now (only with wait: false and no callback_url)
JOB_TIMEOUT504YesThe hold ran out and the job continues — or the job's deadline passed
JOB_FAILED502MaybeEvery attempt failed. The code may be more specific, such as PROMPT_REJECTED
JOB_CANCELLED499NoThe job was cancelled while you waited

Full handling guidance: Errors.