Guides
Rate limits
Every answer carries three headers saying what your budget is and when it resets. A refusal carries them too, so a client never has to discover a limit by hitting it.
The three headers
HTTP/1.1 200 OK
Cache-Control: private, no-store
RateLimit-Limit: 600
RateLimit-Remaining: 599
RateLimit-Reset: 60| Header | Value | Meaning |
|---|---|---|
RateLimit-Limit | integer | The ceiling for this call, after the operation's rate class and any per key limit have both been applied. The lower of the two always wins. |
RateLimit-Remaining | integer | How many calls are left in the current window. Never negative. |
RateLimit-Reset | integer seconds | How long until the window reopens. The window is fixed and anchored at the first request in it, not sliding. |
They describe the limit that applied to the call. On a 429 raised by the account ceiling they describe the account's window, not the key's, because that is the window the caller is waiting for.
They are on refusals too
A scope failure, an account state failure and a validation failure are all refused before they count against any limit, so they have not spent any budget. The trio is still attached, without spending anything, so a caller learns its budget on every answer rather than only on the ones that got through.
HTTP/1.1 403 Forbidden
RateLimit-Limit: 120
RateLimit-Remaining: 120
RateLimit-Reset: 60
{
"error": {
"code": "forbidden_scope",
"message": "This credential does not carry the scope this operation needs.",
"details": { "requiredScopes": ["keys:manage"], "missingScopes": ["keys:manage"] },
"requestId": "68cd298d-ad08-4b8a-a5a0-e0a87fd980b7"
}
}The 4 classes
The class is a property of the operation, so the cost of a call is decided by what it does. Each class is its own budget: spending your whole write budget leaves the read budget untouched.
| Class | Per minute, per credential | What is in it |
|---|---|---|
read | 600 | Cheap reads. Every list and every get. |
write | 120 | Ordinary writes: cancels, saved templates, webhook endpoints, the brand kit. |
expensive | 20 | Heavier work on every request: planning a storyboard, rewriting or suggesting a brief, scanning a website, generating a marketing project, a sentence edit, a voice preview, an assistant message. |
charge | 20 | Operations that spend the account's seconds: starting a render, cutting a marketing export. |
Two of those are not obvious from the verb. renders.quote is a POST and is in the read class, because it prices options and creates nothing. marketing.generations.create is in the expensive class rather than charge, because generating is heavy work and reserves no video time; the export is what spends.
The three ceilings
A call is checked against up to three buckets in order. The first one it is over is the one it is told about, and details.scope in the 429 body says which.
| Order | Bucket | details.scope | Number |
|---|---|---|---|
| 1 | The credential, per rate class, per minute. A key has its own bucket; a signed in person shares one per account. Never per IP address. | minute | The class limit, or the key's own rateLimitPerMinute when it is lower. |
| 2 | The account, across every key and every class, per minute. An account may hold 25 keys, so without this its real throughput would be keys multiplied by classes. | account | 2000 by default. |
| 3 | The key's own daily cap, in UTC days. Only when the key carries one, and never for a signed in person. | day | The key's rateLimitPerDay. |
The credential limit is checked first: a caller already over its own key's budget is told about the smaller and more actionable limit, and a call refused there does not also spend the account's. The daily cap counts requests the same way the usage page does.
Callers with no credential
A request whose credential cannot be resolved is counted per IP address, at 60 a minute. Only failures are counted, so a working integration on the same office network as a script guessing keys is not limited by the guesser's failures.
Over the line, the 401 becomes a 429. The trio is on the 401s as well, so a client can see it is being counted before it is cut off.
What a 429 looks like
HTTP/1.1 429 Too Many Requests
RateLimit-Limit: 20
RateLimit-Remaining: 0
RateLimit-Reset: 12
Retry-After: 12
{
"error": {
"code": "rate_limited",
"message": "Too many requests. Try again shortly.",
"details": {
"class": "charge",
"scope": "minute",
"limit": 20,
"remaining": 0,
"resetAt": "2026-09-21T16:18:17.242Z"
},
"requestId": "247a5860-44ea-437b-8e62-421a2bfda458"
}
}| Field | What it is |
|---|---|
details.class | The rate class of the operation: read, write, expensive, charge, or unauthenticated. |
details.scope | Which ceiling was hit: minute, account or day. A voice preview can also answer day for its own daily cap, or global when the product-wide daily cap on previews is spent. |
details.limit | The number that was exceeded. |
details.resetAt | An ISO timestamp for when the window reopens. The same number as Retry-After, absolute instead of relative. |
Retry-After | Whole seconds. Read this, not the body, unless you need an absolute time. |
Backing off correctly
Three rules, and the code below is all three.
- Obey
Retry-After. The server knows how long the window has left. Exponential backoff is what you use when it does not say, which on this surface means never for a 429. - Jitter. A fleet refused in the same second comes back in the same second unless you add some.
- Retry 429 and 503 only. Every other 4xx is refused for a reason that a second identical request will meet again.
const BASE = "https://api.gogoscreen.com/api/v1";
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
/**
* One call, with backoff that obeys Retry-After.
*
* Retries only 429 and 503. Never retries a 4xx that is not 429: the same
* request will be refused the same way, and on a spending operation a retry
* with a fresh idempotency key is how a video gets paid for twice.
*/
async function call(path, init = {}, { attempts = 5 } = {}) {
for (let attempt = 1; ; attempt += 1) {
const res = await fetch(BASE + path, {
...init,
headers: {
Authorization: `Bearer ${process.env.GOGOSCREEN_API_KEY}`,
...(init.body ? { "Content-Type": "application/json" } : {}),
...init.headers,
},
});
if (res.status !== 429 && res.status !== 503) return res;
if (attempt >= attempts) return res;
// The server says how long. Only guess when it did not.
const header = Number(res.headers.get("retry-after"));
const wait = Number.isFinite(header) && header > 0
? header * 1000
: Math.min(30_000, 2 ** attempt * 500);
// Full jitter on top, so a fleet that was refused together does not come
// back together.
await sleep(wait + Math.random() * 1000);
}
}
/** Slow down before you are refused: the trio is on every answer. */
function paceFrom(res) {
const remaining = Number(res.headers.get("ratelimit-remaining"));
const reset = Number(res.headers.get("ratelimit-reset"));
if (!Number.isFinite(remaining) || !Number.isFinite(reset)) return 0;
if (remaining > 10) return 0;
// Spread what is left over the rest of the window.
return Math.ceil((reset * 1000) / Math.max(1, remaining));
}