Skip to content

Revoye Cloud API Documentation

Updated
Reading time
2 min

Revoye Cloud is a hosted API that gives you ChatGPT, Claude, Gemini, DeepSeek, Qwen and Perplexity through one endpoint and one API key. You send a prompt to https://cloudapi.revoye.com/v1/completions with a bearer token, Revoye Cloud processes it as a job on the model you chose (or on any available model), and you get the answer back — on the same connection, by polling, or by webhook. There is nothing to install, nothing to connect, and no provider account or provider API key to bring.

Before the API will do anything

Two things:

  1. A Revoye Cloud account with a verified email address — create one.
  2. An API key, created in the dashboard.

That is all. The quickstart takes about five minutes.

Start here

QuickstartKey, status check, first prompt. About five minutes
AuthenticationCreating keys, scopes, rotation, what a key cannot do
Node.js integration skillOne downloadable file covering the whole public API. Hand it to an AI coding agent, or read it yourself

API reference

POST /v1/completionsThe native endpoint. Everything the API can do is expressible here
POST /v1/chat/completionsOpenAI-shaped, for existing SDKs
GET /v1/modelsThe models you can use right now, in the shape a model picker expects
GET /v1/statusAvailable capacity per model, hourly usage, and your queue
OpenAPI specificationThe machine-readable contract

Guides

Retrying safelyIdempotency keys, and why a lost response is not a lost job
WebhooksReceiving finished jobs, and at-least-once delivery
ErrorsEvery code, whether to retry, and which ones mean "wait" rather than "failed"
Rate limits and capacityRequest limits, hourly model limits, and the difference
Using an OpenAI SDKOne base URL change, and the things that differ
Best practicesDesigning an integration around job-based processing

Three things to know before you build

Every request is a job. Revoye Cloud records the job the moment it is accepted, and gives it an id. wait: true is a long poll on that job, not a synchronous call: a dropped connection loses nothing — the result is at GET /v1/completions/{id} or on its way to your webhook.

A response typically takes 20 to 90 seconds. Set long HTTP timeouts, or do not hold a connection at all: queue work, take webhooks, and do not put the API behind a spinner someone is watching.

Capacity per model is finite. When every worker for a model is busy, or the model's hourly limit is used up, your job waits in the queue until it can run. That is a feature for batch work and a constraint for everything else.

Everything is versioned under /v1

/v1 is a stable contract. Additive changes — a new optional field, a new error code — may ship at any time. Anything that would break a working integration ships as /v2 alongside it.