Skip to content

POST /v1/chat/completions

Updated
Reading time
3 min

An OpenAI-shaped front door to the same jobs as /v1/completions, so a client built for the OpenAI chat format can submit work to Revoye Cloud by changing one base URL. It translates the request in, creates a job exactly as the native endpoint does, and answers in the chat-completion envelope. Requires the completions:write scope.

{
  "model": "revoye/chatgpt",
  "messages": [
    { "role": "system", "content": "You are terse." },
    { "role": "user", "content": "Summarise the CAP theorem." }
  ],
  "stream": false
}

Model names

modelRuns on
revoye/autoAny available model
revoye/chatgptChatGPT
revoye/claudeClaude
revoye/geminiGemini
revoye/deepseekDeepSeek
revoye/qwenQwen
revoye/perplexityPerplexity

The revoye/ prefix is optional: chatgpt and auto are accepted too, which is what lets an id from GET /v1/models be sent back as model unchanged. Anything else is 400 INVALID_REQUEST with details.field: "model".

The response is the job, not yet the answer

The endpoint returns 202 as soon as the job is accepted, without waiting for it to finish:

{
  "id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9",
  "object": "chat.completion",
  "created": 1758186842,
  "model": "revoye/chatgpt",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": null },
      "finish_reason": null
    }
  ],
  "usage": null,
  "_revoye": { "job_id": "job_01JAY7Q2K8XYZ3M4N5P6Q7R8S9", "status": "queued" }
}

choices[0].message.content is null until the job has succeeded. Collect the answer one of two ways:

  • Poll the job. GET /v1/completions/{id} with the _revoye.job_id returns the native job object; its response field holds the answer once status is succeeded. This needs the completions:read scope.
  • Repeat the request with the same Idempotency-Key. An identical request carrying the same key returns the same job rather than creating a new one, and once it has succeeded the envelope comes back with content filled in and finish_reason: "stop". Use a fresh key for every new prompt.

A response typically takes 20 to 90 seconds, so poll every few seconds rather than continuously.

What else behaves differently

messages are flattened into one prompt. Every system message comes first, then the other turns in order as User: … and Assistant: …, separated by blank lines. This is not a stateful conversation API — to continue a real conversation, use the native endpoint with conversation_ref and mode: "continue".

Because they become one prompt, they share one prompt's limits: at most 256 messages, each at most 100 000 characters, and the flattened result at most 100 000 characters. Over the first two is 400 INVALID_REQUEST; over the flattened total is 413 PAYLOAD_TOO_LARGE.

stream: true returns 400. Revoye Cloud returns complete answers only. Silently ignoring the flag would look like a hang, which is worse than an error.

usage is null. The API does not report token counts. If your code does arithmetic on usage, guard it.

Other OpenAI parameters are ignored. temperature, max_tokens, tools and the like are not read; only model, messages and stream are.

When to use the native endpoint instead

The compatibility layer exists to save you a rewrite. Everything else is on /v1/completions: wait: true long polling, callback_url webhooks, timeout_ms, deadline_ms, priority, metadata, and conversation continuation. For new code, start there.

Errors

Same envelope and same codes as /v1/completions, plus INVALID_REQUEST for stream: true and for an unrecognised model. See Errors.