Revoye Cloud API Documentation
- Updated
- Reading time
- 2 min
Revoye Cloud is a hosted API that gives you ChatGPT, Claude, Gemini, DeepSeek, Qwen and
Perplexity through one endpoint and one API key. You send a prompt to
https://cloudapi.revoye.com/v1/completions with a bearer token, Revoye Cloud processes it as a job on the model
you chose (or on any available model), and you get the answer back — on the same connection, by
polling, or by webhook. There is nothing to install, nothing to connect, and no provider account or
provider API key to bring.
Before the API will do anything
Two things:
- A Revoye Cloud account with a verified email address — create one.
- An API key, created in the dashboard.
That is all. The quickstart takes about five minutes.
Start here
| Quickstart | Key, status check, first prompt. About five minutes |
| Authentication | Creating keys, scopes, rotation, what a key cannot do |
| Node.js integration skill | One downloadable file covering the whole public API. Hand it to an AI coding agent, or read it yourself |
API reference
POST /v1/completions | The native endpoint. Everything the API can do is expressible here |
POST /v1/chat/completions | OpenAI-shaped, for existing SDKs |
GET /v1/models | The models you can use right now, in the shape a model picker expects |
GET /v1/status | Available capacity per model, hourly usage, and your queue |
| OpenAPI specification | The machine-readable contract |
Guides
| Retrying safely | Idempotency keys, and why a lost response is not a lost job |
| Webhooks | Receiving finished jobs, and at-least-once delivery |
| Errors | Every code, whether to retry, and which ones mean "wait" rather than "failed" |
| Rate limits and capacity | Request limits, hourly model limits, and the difference |
| Using an OpenAI SDK | One base URL change, and the things that differ |
| Best practices | Designing an integration around job-based processing |
Three things to know before you build
Every request is a job. Revoye Cloud records the job the moment it is accepted, and gives it an
id. wait: true is a long poll on that job, not a synchronous call: a dropped connection loses
nothing — the result is at GET /v1/completions/{id} or on its way to your webhook.
A response typically takes 20 to 90 seconds. Set long HTTP timeouts, or do not hold a connection at all: queue work, take webhooks, and do not put the API behind a spinner someone is watching.
Capacity per model is finite. When every worker for a model is busy, or the model's hourly limit is used up, your job waits in the queue until it can run. That is a feature for batch work and a constraint for everything else.
Everything is versioned under /v1
/v1 is a stable contract. Additive changes — a new optional field, a new error code — may ship at
any time. Anything that would break a working integration ships as /v2 alongside it.