POST /v1/chat/completions
- Updated
- Reading time
- 3 min
An OpenAI-shaped front door to the same jobs as /v1/completions, so a
client built for the OpenAI chat format can submit work to Revoye Cloud by changing one base URL. It
translates the request in, creates a job exactly as the native endpoint does, and answers in the
chat-completion envelope. Requires the completions:write scope.
{
"model": "revoye/chatgpt",
"messages": [
{ "role": "system", "content": "You are terse." },
{ "role": "user", "content": "Summarise the CAP theorem." }
],
"stream": false
}Model names
model | Runs on |
|---|---|
revoye/auto | Any available model |
revoye/chatgpt | ChatGPT |
revoye/claude | Claude |
revoye/gemini | Gemini |
revoye/deepseek | DeepSeek |
revoye/qwen | Qwen |
revoye/perplexity | Perplexity |
The revoye/ prefix is optional: chatgpt and auto are accepted too, which is what lets an id
from GET /v1/models be sent back as model unchanged. Anything else is
400 INVALID_REQUEST with details.field: "model".
The response is the job, not yet the answer
The endpoint returns 202 as soon as the job is accepted, without waiting for it to finish:
{
"id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9",
"object": "chat.completion",
"created": 1758186842,
"model": "revoye/chatgpt",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": null },
"finish_reason": null
}
],
"usage": null,
"_revoye": { "job_id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9", "status": "queued" }
}choices[0].message.content is null until the job has succeeded. Collect the answer one of two
ways:
- Poll the job.
GET /v1/completions/{id}with the_revoye.job_idreturns the native job object; itsresponsefield holds the answer oncestatusissucceeded. This needs thecompletions:readscope. - Repeat the request with the same
Idempotency-Key. An identical request carrying the same key returns the same job rather than creating a new one, and once it has succeeded the envelope comes back withcontentfilled in andfinish_reason: "stop". Use a fresh key for every new prompt.
A response typically takes 20 to 90 seconds, so poll every few seconds rather than continuously.
What else behaves differently
messages are flattened into one prompt. Every system message comes first, then the other
turns in order as User: … and Assistant: …, separated by blank lines. This is not a stateful
conversation API — to continue a real conversation, use the native endpoint with conversation_ref
and mode: "continue".
Because they become one prompt, they share one prompt's limits: at most 256 messages, each at most
100 000 characters, and the flattened result at most 100 000 characters. Over the first two is
400 INVALID_REQUEST; over the flattened total is 413 PAYLOAD_TOO_LARGE.
stream: true returns 400. Revoye Cloud returns complete answers only. Silently ignoring the
flag would look like a hang, which is worse than an error.
usage is null. The API does not report token counts. If your code does arithmetic on
usage, guard it.
Other OpenAI parameters are ignored. temperature, max_tokens, tools and the like are not
read; only model, messages and stream are.
When to use the native endpoint instead
The compatibility layer exists to save you a rewrite. Everything else is on
/v1/completions: wait: true long polling, callback_url webhooks,
timeout_ms, deadline_ms, priority, metadata, and conversation continuation. For new code,
start there.
Errors
Same envelope and same codes as /v1/completions, plus INVALID_REQUEST
for stream: true and for an unrecognised model. See Errors.