Skip to content

Using an OpenAI SDK with Revoye Cloud

Updated
Reading time
3 min
Level
beginner

Set the base URL to https://cloudapi.revoye.com/v1, use your Revoye Cloud key as the API key, and set the model to revoye/auto or revoye/<provider>. The call returns as soon as the job is accepted, with content: null; repeat the same call with the same Idempotency-Key until the content arrives. The examples below wrap that in a small helper.

Python

import os, time, uuid
from openai import OpenAI
 
client = OpenAI(
    base_url=f"{os.environ['REVOYE_BASE_URL']}/v1",
    api_key=os.environ["REVOYE_API_KEY"],
)
 
def ask(messages, model="revoye/auto", poll_seconds=5, give_up_after=900):
    key = str(uuid.uuid4())                       # one key per prompt, reused while polling
    deadline = time.monotonic() + give_up_after
    while True:
        answer = client.chat.completions.create(
            model=model,
            messages=messages,
            extra_headers={"Idempotency-Key": key},
        )
        content = answer.choices[0].message.content
        if content is not None:
            return content
        status = answer.model_extra["_revoye"]["status"]
        if status in ("failed", "cancelled", "expired") or time.monotonic() > deadline:
            raise RuntimeError(f"job {answer.id} ended as {status}")
        time.sleep(poll_seconds)
 
print(ask([{"role": "user", "content": "Summarise the CAP theorem."}], model="revoye/chatgpt"))

TypeScript

import OpenAI from "openai";
 
const client = new OpenAI({
  baseURL: `${process.env.REVOYE_BASE_URL}/v1`,
  apiKey: process.env.REVOYE_API_KEY,
});
 
type Message = { role: "system" | "user" | "assistant"; content: string };
 
export async function ask(messages: Message[], model = "revoye/auto"): Promise<string> {
  const key = crypto.randomUUID();                // one key per prompt, reused while polling
  for (;;) {
    const answer = await client.chat.completions.create(
      { model, messages },
      { headers: { "Idempotency-Key": key } },
    );
    const content = answer.choices[0]?.message.content;
    if (content != null) return content;
 
    const status = (answer as unknown as { _revoye: { status: string } })._revoye.status;
    if (["failed", "cancelled", "expired"].includes(status)) {
      throw new Error(`job ${answer.id} ended as ${status}`);
    }
    await new Promise((resolve) => setTimeout(resolve, 5_000));
  }
}

Repeating the request is safe because the key is the same: Revoye Cloud returns the existing job each time rather than creating a new one. Never set Idempotency-Key as a client-wide default header — every prompt would then map to the first job. If you would rather not repeat the request, read _revoye.job_id from the first response and poll GET /v1/completions/{id} instead.

What differs from a model API

The first response has no answer. choices[0].message.content is null and finish_reason is null until the job has succeeded; a response typically takes 20 to 90 seconds. Code that reads content from the first response will see null.

Messages are flattened. Every system message is placed first and the remaining turns follow as User: … and Assistant: … in a single prompt. If you need to continue an actual conversation, use the native endpoint with conversation_ref and mode: "continue".

stream: true returns 400. There is nothing to stream. Any UI code that assumes a stream needs a non-streaming path before it can talk to Revoye Cloud.

usage is null. The API does not report token counts. Guard any code that does arithmetic on it — cost estimation, budget enforcement, context-window maths — because None + int is the failure you will hit first.

Only model, messages and stream are read. Sampling parameters such as temperature and max_tokens, and tool definitions, are accepted and ignored.

Model pickers

GET /v1/models returns the enabled models in OpenAI's list shape, with revoye/auto first, so client.models.list() populates a picker with no changes. Every id it returns is accepted as model.

What you give up

The compatibility layer covers submitting a prompt and collecting the answer. Everything else lives on /v1/completions: long-poll wait: true, webhooks, per-attempt timeouts, job deadlines, priority, metadata, and continuing a conversation. For new code, use the native endpoint. The compatibility layer is for saving a rewrite, not for building on.