Developers/API Reference/Rate Limits & Quotas
Rate Limits & Quotas
How Supero limits traffic: the per-surface request-rate limits and their 429 contract, the plan quota system and its 402 contract, and the per-request size and count limits that apply to every caller.
Overview
Two different mechanisms can refuse a call, and they behave nothing alike. Confusing them is the usual reason a client retries something it should not, or gives up on something it should retry.
| Rate limits | Quotas | |
|---|---|---|
| Answers | Too many requests, too quickly | You have used up what your plan includes |
| Status | 429 | 402 |
| Window | Seconds to an hour | A billing period |
| Fix | Wait and retry | Upgrade, buy credit, or wait for the period to roll over |
| Retrying helps | Yes, after the stated delay | No — the same call fails until something changes |
On top of both there is a third category: fixed per-request limits — how big a body can be, how many records come back in a page, how many files a bundle may contain. Those are the same for everyone regardless of plan, and they are the most useful numbers on this page because you can design against them.
Request Rate Limits
Supero does not publish a single global requests-per-second figure for an account. Rate limiting is applied per surface, on the endpoints where abuse or cost makes it necessary, and each one has its own window and its own key.
| Surface | Limit | Counted per | Response |
|---|---|---|---|
| Public (unauthenticated) reads under /api/v1/public/... | 100 requests / minute | Project and caller IP | 429 with Retry-After: 60 |
| Public image serving | 600 requests / minute | Caller IP | 429 with Retry-After: 60 |
| App generation (POST /ai/v1/app/generate, /ai/v1/app/upload-bundle) | 10 requests / minute, and 2 generations running at once | Domain and caller IP | 429. Concurrent generations queue rather than fail. |
| Deploy log reads | A short cooldown between reads | Domain | 429 with Retry-After, plus a retry delay in the body |
| Log ingest | 30 requests / minute | Domain | 429 with a retry delay in the body |
| Domain registration | A per-IP ceiling | Caller IP | 429 |
| Verification-code emails (signup) | 3 sends / 10 minutes | Email address | 429 |
| Password-reset emails | 5 sends / hour | Email address | 429 |
| SDK generation | A cooldown between builds | Build request | 429 with Retry-After |
| Connector operations | Set by the connector, and by the source it talks to | Connector | 429 with Retry-After and a retry_after field in the body |
ℹ️ Ordinary authenticated CRUD is not rate limited per-request today
Reads and writes against /api/v1/crud/... are not subject to a published per-second or per-minute limit. What bounds them is your plan's monthly API-request allowance and the per-request limits further down this page. This is not a promise that no limit will ever apply — build a client that handles a 429 anywhere, because a limit can be introduced on any surface.
⚠️ Treat the figures above as a floor, not a threshold
These are the limits the platform enforces, not a target to pace against. A client tuned to sit just under a published number will trip it. Leave headroom, and back off on 429 rather than assuming the limit was higher than documented.
A request-rate ceiling also applies at the network edge, ahead of any application-level limit, sized for normal client traffic. It is not published as a contract and is not something a well-behaved client will meet. If you have a legitimate high-volume workload — a bulk migration, a nightly sync, a load test — tell us before you run it: [email protected].
The 429 Contract
A rate-limited call returns 429 with a JSON body. The public API family is the one most likely to hit a limit, and it returns:
json
{
"success": false,
"error": "Rate limit exceeded. Please try again later."
}| Header | Present? | Meaning |
|---|---|---|
| Retry-After | On public API, public image, and SDK-build 429s | Seconds to wait before retrying. For the public API family the value is 60. |
| X-RateLimit-Limit / X-RateLimit-Remaining / X-RateLimit-Reset | Not part of the public contract | Do not build a client that reads a remaining-quota header from api.supero.dev — no successful response carries one. |
Because there is no remaining-count header to read, a client cannot pace itself by inspecting responses. Handle rate limiting reactively instead:
- •Retry on 429, and only on 429 — a 402 will never succeed on retry.
- •Honour Retry-After when it is present. When it is absent, start at a few seconds and back off exponentially with jitter.
- •Cap total retries. A limit measured per minute clears in under a minute; something still failing after several minutes is a different problem.
- •Never retry in a tight loop. Retrying immediately is what turns a brief limit into a sustained one.
💡 The SDK surfaces this as a typed error
The Python client raises RateLimitError, carrying retry_after, so you can catch it separately from AuthenticationError and AuthorizationError rather than string-matching a message.
Quotas and the 402 Contract
A quota is a plan allowance measured over a billing period. When an operation would exceed one, the call returns 402 with a structured body naming exactly which meter tripped, so a client can show a real message instead of a generic failure.
Two body shapes exist, and the difference is only where the object sits. Control-plane and CRUD endpoints return the object at the top level:
json
{
"error": "quota_exceeded",
"upgrade_required": true,
"meter": "projects",
"limit": 3,
"used": 3,
"remaining": 0,
"policy": "alert_then_block",
"reason": "Project limit reached for your plan",
"message": "Project limit reached for your plan"
}AI endpoints under /ai/v1 return the same object nested under detail:
json
{
"detail": {
"error": "quota_exceeded",
"upgrade_required": true,
"meter": "agent.prompt",
"limit": 10,
"used": 10,
"remaining": 0,
"policy": "block",
"reason": "10 free AI queries this month used",
"message": "10 free AI queries this month used"
}
}| Field | Type | Meaning |
|---|---|---|
| error | string | Always quota_exceeded. Branch on this rather than parsing message. |
| upgrade_required | boolean | Always true on a block. A hint for the UI, not a separate condition. |
| meter | string | Which allowance was exhausted — see the meter list below. |
| limit | integer | The plan's allowance for that meter. A negative value or null means unlimited. |
| used | number | Consumption so far in the current period. |
| remaining | number | What is left. Zero or negative on a block. |
| policy | string | The overage policy in force: block, alert_then_block, or alert_only. |
| reason | string | A human-readable explanation of the specific block. |
| message | string | A display string, falling back to a generic upgrade prompt when reason is empty. |
⚠️ Not every 402 carries every field
Some plan gates return a shorter version with only error, upgrade_required, meter, reason and message — no limit, used or remaining. A marketplace entitlement block is shorter still, carrying only the two-field error shape with PAYMENT_REQUIRED in the text. Read the fields defensively: branch on the status and on error, and treat the numeric fields as optional.
⚠️ Do not retry a 402
A 402 is a state, not a transient failure. The same call will fail identically until the plan changes, credit is added, or the billing period rolls over. Surface it to a human — the body has everything an upgrade prompt needs — and stop.
What Triggers a 402
Quota checks run on the operations that consume real capacity, not on every request. These are the ones that will stop you:
| Operation | Meter | Typical cause |
|---|---|---|
| Publishing a live deployment | servers | The plan includes no live deployments, or all of them are in use. The lower tiers are build-and-preview only. |
| Creating a project | projects | At the plan's project limit. |
| Generating an app | appgen.run | At the plan's monthly app-generation allowance. |
| Sending an AI-agent prompt | agent.prompt | Past the free monthly allowance without the AI Agent add-on or a plan that bundles it. |
| Launching a preview, running a connector, deploying | power | The monthly Power allowance is exhausted and no prepaid credit covers the shortfall. |
Other plan dimensions — monthly API requests, storage, tenant count, team members — are metered and reported through usage rather than refused at request time. That means a 402 is not a reliable signal that you are approaching those ceilings. Read your usage directly and alert on it yourself; see Billing & Usage API.
ℹ️ Prepaid credit is consulted before you are blocked
For the Power meter specifically, if you are over the allowance but hold enough prepaid Power credit to cover the shortfall, the request is allowed and the difference is drawn from your credit balance instead of returning 402.
Meters
The meter named in a 402 comes from a fixed catalog. Most of them draw on Power — a single usage currency — at different weights, so one Power figure covers activity across the platform instead of a separate counter per action.
| Meter | What it counts | Power per unit |
|---|---|---|
| appgen.run | A full app-generation run | 500 |
| deploy.run | An app deployment | 200 |
| schemagen.run | A schema-generation run | 200 |
| agent.prompt | One AI-agent user turn | 100 |
| testdata.run | A test-data generation run | 100 |
| cloudrun.hours | Hosted runtime hours | 60 per hour |
| connector.execution / connector.run_query | A connector sync run or live warehouse query | 50 |
| mcp.build | An app publish or generation over MCP | 20 |
| compute.workflow | A workflow execution | 10 |
| connector.poll / mcp.tool_call | A connector poll cycle; an MCP tool invocation | 5 |
| external.trigger | An external trigger or webhook invocation | 2 |
| image.fetch | An external image-provider fetch | 1 |
| api.request | An API request | 0.01 |
Storage is metered separately from Power, on its own card, against your plan's storage allowance. A further set of meters — token counts, request reads and writes, egress bytes, record counts, plugin compute time, and point-in-time gauges for projects, tenants and seats — is recorded for reporting and does not draw Power.
For the allowances each plan grants across these dimensions, see Plans & Billing.
Per-Request Limits
These apply to every caller on every plan. They are fixed, they are enforced, and they are the numbers to design against.
Request size
| Limit | Value | On exceeding |
|---|---|---|
| File upload, single file | 25 MB | Rejected |
| File upload, files per batch | 10 | Rejected |
| MCP JSON-RPC request body | 12 MB | 413, with a JSON-RPC error telling you to send fewer or smaller files |
| App bundle, files | 300 | Rejected with a message naming the count and the limit |
| App bundle, bytes per file | 2 MB | Rejected |
| App bundle, total size | 8 MB | Rejected |
A large body may also be refused at the network edge before the application sees it. If a request near any of these ceilings fails without a JSON body, assume it was rejected on size and split it.
Pagination
| Limit | Value | Behaviour |
|---|---|---|
| Maximum page size | 1000 | A larger limit is silently clamped to 1000 — no error is returned, so check how many rows you actually received. |
| Default page size | 50 | Applied only once at least one pagination parameter is present. |
| Public (unauthenticated) reads, maximum page size | 100 | A tighter cap than the authenticated API. |
🚨 No pagination parameters means no page size at all
Send no pagination parameter and the read is unbounded: every matching record comes back and the response carries no pagination object. The default of 50 engages only when you send some pagination parameter but omit limit — so ?sort_by=name alone silently caps you at 50 rows, while sending nothing at all returns the whole collection. Always send an explicit limit. See Querying & Pagination.
Batch writes
| Limit | Value |
|---|---|
| Errors accumulated before the batch stops processing | 100 |
| Errors returned in the response | The first 10 |
| Length of each error message | Truncated to 200 characters |
A batch is not rejected for being large, but 1000 records is the recommended ceiling — beyond that, split the call. Note that the response can stop short of your input: once 100 errors accumulate the remaining records are not attempted, and the error list you get back covers only the first 10. Reconcile using the created, updated and failed counts, not the error list. A batch that partially fails still returns HTTP 200 — see Errors & Status Codes.
AI and workflow bounds
| Limit | Value | On exceeding |
|---|---|---|
| Chat message length | 1 to 10,000 characters | 422 |
| Chat agent iterations | 1 to 10, default 5 | 422 |
| App-generation description | 10 to 12,000 characters | 422 |
| Vector document content | 1 to 50,000 characters | 422 |
| Vector batch index size | 1 to 100 documents | 422 |
| Semantic search query | 1 to 2,000 characters; top_k 1 to 50 | 422 |
| Workflow loop iterations | Default ceiling 10 per step, and never more than 100 | Silently clamped to the ceiling |
| Workflow step timeout | 30 seconds by default | The step fails on timeout and its on_error applies |
Sessions
| Credential | Lifetime |
|---|---|
| Access token | 8 hours (expires_in: 28800) |
| Refresh token | 7 days |
| Verification and password-reset codes | 10 minutes |
What This Page Does Not Tell You
Being straight about the boundaries of this page is more useful than a table that looks complete:
- •There is no single published requests-per-second or requests-per-minute figure for an account or an API key. Limits are per surface, and the surfaces that carry one are listed above.
- •An API key can record an hourly request rate when it is created. That value is stored with the key but is not applied at the public API edge today, so it is not a control you can rely on — do not use it as a way to throttle a key.
- •The per-minute figures above can change with deployment sizing. They are what the platform enforces, not a contractual floor.
- •Self-hosted and on-premise deployments meter differently. Plan allowances are a hosted-platform concept; a self-hosted install is governed by its licence.
ℹ️ Getting the numbers that apply to you
For the limits in force on your own domain, plan, or deployment — and for raising any of them ahead of a launch, a migration, or a load test — contact [email protected]. Include your domain name, the endpoint, and the rate you need. Current plan allowances are readable at any time from the Billing screen in the console, or from GET /api/v1/billing/usage/summary.
Designing for Limits
- •Send an explicit limit on every list call, and page with a cursor rather than a deep offset.
- •Prefer one filtered POST /query over pulling a collection and filtering client-side — the filtered call is one request against your allowance instead of many.
- •Use the aggregate endpoints for counts and sums instead of fetching rows to count them.
- •Batch writes, but keep batches modest and check the returned counts rather than the status code.
- •Cache reads that do not change often. Public reads are already cached briefly, so a tight polling loop mostly re-reads the same cached answer while still consuming your allowance.
- •Handle 429 with backoff and 402 by surfacing it, and never conflate the two in one retry path.
- •Read your usage on a schedule and alert before you hit a ceiling, rather than discovering it as a failed request on launch day.
Next Steps
- •Errors & Status Codes — the full status-code catalogue, including the failures that return 200.
- •Billing & Usage API — reading your usage and allowances programmatically.
- •Plans & Billing — the plan catalog, Power, credit packs, and overage handling.
- •Querying & Pagination — page sizes, cursors, and the filter grammar.
- •Going to Production — checking your plan before launch day.
On this page