S
supero.docs
Developers/API Reference/Rate Limits & Quotas

Rate Limits & Quotas

How Supero limits traffic: the per-surface request-rate limits and their 429 contract, the plan quota system and its 402 contract, and the per-request size and count limits that apply to every caller.

Overview

Two different mechanisms can refuse a call, and they behave nothing alike. Confusing them is the usual reason a client retries something it should not, or gives up on something it should retry.
Rate limitsQuotas
AnswersToo many requests, too quicklyYou have used up what your plan includes
Status429402
WindowSeconds to an hourA billing period
FixWait and retryUpgrade, buy credit, or wait for the period to roll over
Retrying helpsYes, after the stated delayNo — the same call fails until something changes
On top of both there is a third category: fixed per-request limits — how big a body can be, how many records come back in a page, how many files a bundle may contain. Those are the same for everyone regardless of plan, and they are the most useful numbers on this page because you can design against them.

Request Rate Limits

Supero does not publish a single global requests-per-second figure for an account. Rate limiting is applied per surface, on the endpoints where abuse or cost makes it necessary, and each one has its own window and its own key.
SurfaceLimitCounted perResponse
Public (unauthenticated) reads under /api/v1/public/...100 requests / minuteProject and caller IP429 with Retry-After: 60
Public image serving600 requests / minuteCaller IP429 with Retry-After: 60
App generation (POST /ai/v1/app/generate, /ai/v1/app/upload-bundle)10 requests / minute, and 2 generations running at onceDomain and caller IP429. Concurrent generations queue rather than fail.
Deploy log readsA short cooldown between readsDomain429 with Retry-After, plus a retry delay in the body
Log ingest30 requests / minuteDomain429 with a retry delay in the body
Domain registrationA per-IP ceilingCaller IP429
Verification-code emails (signup)3 sends / 10 minutesEmail address429
Password-reset emails5 sends / hourEmail address429
SDK generationA cooldown between buildsBuild request429 with Retry-After
Connector operationsSet by the connector, and by the source it talks toConnector429 with Retry-After and a retry_after field in the body

ℹ️ Ordinary authenticated CRUD is not rate limited per-request today

Reads and writes against /api/v1/crud/... are not subject to a published per-second or per-minute limit. What bounds them is your plan's monthly API-request allowance and the per-request limits further down this page. This is not a promise that no limit will ever apply — build a client that handles a 429 anywhere, because a limit can be introduced on any surface.

⚠️ Treat the figures above as a floor, not a threshold

These are the limits the platform enforces, not a target to pace against. A client tuned to sit just under a published number will trip it. Leave headroom, and back off on 429 rather than assuming the limit was higher than documented.
A request-rate ceiling also applies at the network edge, ahead of any application-level limit, sized for normal client traffic. It is not published as a contract and is not something a well-behaved client will meet. If you have a legitimate high-volume workload — a bulk migration, a nightly sync, a load test — tell us before you run it: [email protected].

The 429 Contract

A rate-limited call returns 429 with a JSON body. The public API family is the one most likely to hit a limit, and it returns:
json
{
  "success": false,
  "error": "Rate limit exceeded. Please try again later."
}
HeaderPresent?Meaning
Retry-AfterOn public API, public image, and SDK-build 429sSeconds to wait before retrying. For the public API family the value is 60.
X-RateLimit-Limit / X-RateLimit-Remaining / X-RateLimit-ResetNot part of the public contractDo not build a client that reads a remaining-quota header from api.supero.dev — no successful response carries one.
Because there is no remaining-count header to read, a client cannot pace itself by inspecting responses. Handle rate limiting reactively instead:
  • •Retry on 429, and only on 429 — a 402 will never succeed on retry.
  • •Honour Retry-After when it is present. When it is absent, start at a few seconds and back off exponentially with jitter.
  • •Cap total retries. A limit measured per minute clears in under a minute; something still failing after several minutes is a different problem.
  • •Never retry in a tight loop. Retrying immediately is what turns a brief limit into a sustained one.

💡 The SDK surfaces this as a typed error

The Python client raises RateLimitError, carrying retry_after, so you can catch it separately from AuthenticationError and AuthorizationError rather than string-matching a message.

Quotas and the 402 Contract

A quota is a plan allowance measured over a billing period. When an operation would exceed one, the call returns 402 with a structured body naming exactly which meter tripped, so a client can show a real message instead of a generic failure.
Two body shapes exist, and the difference is only where the object sits. Control-plane and CRUD endpoints return the object at the top level:
json
{
  "error": "quota_exceeded",
  "upgrade_required": true,
  "meter": "projects",
  "limit": 3,
  "used": 3,
  "remaining": 0,
  "policy": "alert_then_block",
  "reason": "Project limit reached for your plan",
  "message": "Project limit reached for your plan"
}
AI endpoints under /ai/v1 return the same object nested under detail:
json
{
  "detail": {
    "error": "quota_exceeded",
    "upgrade_required": true,
    "meter": "agent.prompt",
    "limit": 10,
    "used": 10,
    "remaining": 0,
    "policy": "block",
    "reason": "10 free AI queries this month used",
    "message": "10 free AI queries this month used"
  }
}
FieldTypeMeaning
errorstringAlways quota_exceeded. Branch on this rather than parsing message.
upgrade_requiredbooleanAlways true on a block. A hint for the UI, not a separate condition.
meterstringWhich allowance was exhausted — see the meter list below.
limitintegerThe plan's allowance for that meter. A negative value or null means unlimited.
usednumberConsumption so far in the current period.
remainingnumberWhat is left. Zero or negative on a block.
policystringThe overage policy in force: block, alert_then_block, or alert_only.
reasonstringA human-readable explanation of the specific block.
messagestringA display string, falling back to a generic upgrade prompt when reason is empty.

⚠️ Not every 402 carries every field

Some plan gates return a shorter version with only error, upgrade_required, meter, reason and message — no limit, used or remaining. A marketplace entitlement block is shorter still, carrying only the two-field error shape with PAYMENT_REQUIRED in the text. Read the fields defensively: branch on the status and on error, and treat the numeric fields as optional.

⚠️ Do not retry a 402

A 402 is a state, not a transient failure. The same call will fail identically until the plan changes, credit is added, or the billing period rolls over. Surface it to a human — the body has everything an upgrade prompt needs — and stop.

What Triggers a 402

Quota checks run on the operations that consume real capacity, not on every request. These are the ones that will stop you:
OperationMeterTypical cause
Publishing a live deploymentserversThe plan includes no live deployments, or all of them are in use. The lower tiers are build-and-preview only.
Creating a projectprojectsAt the plan's project limit.
Generating an appappgen.runAt the plan's monthly app-generation allowance.
Sending an AI-agent promptagent.promptPast the free monthly allowance without the AI Agent add-on or a plan that bundles it.
Launching a preview, running a connector, deployingpowerThe monthly Power allowance is exhausted and no prepaid credit covers the shortfall.
Other plan dimensions — monthly API requests, storage, tenant count, team members — are metered and reported through usage rather than refused at request time. That means a 402 is not a reliable signal that you are approaching those ceilings. Read your usage directly and alert on it yourself; see Billing & Usage API.

ℹ️ Prepaid credit is consulted before you are blocked

For the Power meter specifically, if you are over the allowance but hold enough prepaid Power credit to cover the shortfall, the request is allowed and the difference is drawn from your credit balance instead of returning 402.

Meters

The meter named in a 402 comes from a fixed catalog. Most of them draw on Power — a single usage currency — at different weights, so one Power figure covers activity across the platform instead of a separate counter per action.
MeterWhat it countsPower per unit
appgen.runA full app-generation run500
deploy.runAn app deployment200
schemagen.runA schema-generation run200
agent.promptOne AI-agent user turn100
testdata.runA test-data generation run100
cloudrun.hoursHosted runtime hours60 per hour
connector.execution / connector.run_queryA connector sync run or live warehouse query50
mcp.buildAn app publish or generation over MCP20
compute.workflowA workflow execution10
connector.poll / mcp.tool_callA connector poll cycle; an MCP tool invocation5
external.triggerAn external trigger or webhook invocation2
image.fetchAn external image-provider fetch1
api.requestAn API request0.01
Storage is metered separately from Power, on its own card, against your plan's storage allowance. A further set of meters — token counts, request reads and writes, egress bytes, record counts, plugin compute time, and point-in-time gauges for projects, tenants and seats — is recorded for reporting and does not draw Power.
For the allowances each plan grants across these dimensions, see Plans & Billing.

Per-Request Limits

These apply to every caller on every plan. They are fixed, they are enforced, and they are the numbers to design against.

Request size

LimitValueOn exceeding
File upload, single file25 MBRejected
File upload, files per batch10Rejected
MCP JSON-RPC request body12 MB413, with a JSON-RPC error telling you to send fewer or smaller files
App bundle, files300Rejected with a message naming the count and the limit
App bundle, bytes per file2 MBRejected
App bundle, total size8 MBRejected
A large body may also be refused at the network edge before the application sees it. If a request near any of these ceilings fails without a JSON body, assume it was rejected on size and split it.

Pagination

LimitValueBehaviour
Maximum page size1000A larger limit is silently clamped to 1000 — no error is returned, so check how many rows you actually received.
Default page size50Applied only once at least one pagination parameter is present.
Public (unauthenticated) reads, maximum page size100A tighter cap than the authenticated API.

🚨 No pagination parameters means no page size at all

Send no pagination parameter and the read is unbounded: every matching record comes back and the response carries no pagination object. The default of 50 engages only when you send some pagination parameter but omit limit — so ?sort_by=name alone silently caps you at 50 rows, while sending nothing at all returns the whole collection. Always send an explicit limit. See Querying & Pagination.

Batch writes

LimitValue
Errors accumulated before the batch stops processing100
Errors returned in the responseThe first 10
Length of each error messageTruncated to 200 characters
A batch is not rejected for being large, but 1000 records is the recommended ceiling — beyond that, split the call. Note that the response can stop short of your input: once 100 errors accumulate the remaining records are not attempted, and the error list you get back covers only the first 10. Reconcile using the created, updated and failed counts, not the error list. A batch that partially fails still returns HTTP 200 — see Errors & Status Codes.

AI and workflow bounds

LimitValueOn exceeding
Chat message length1 to 10,000 characters422
Chat agent iterations1 to 10, default 5422
App-generation description10 to 12,000 characters422
Vector document content1 to 50,000 characters422
Vector batch index size1 to 100 documents422
Semantic search query1 to 2,000 characters; top_k 1 to 50422
Workflow loop iterationsDefault ceiling 10 per step, and never more than 100Silently clamped to the ceiling
Workflow step timeout30 seconds by defaultThe step fails on timeout and its on_error applies

Sessions

CredentialLifetime
Access token8 hours (expires_in: 28800)
Refresh token7 days
Verification and password-reset codes10 minutes

What This Page Does Not Tell You

Being straight about the boundaries of this page is more useful than a table that looks complete:
  • •There is no single published requests-per-second or requests-per-minute figure for an account or an API key. Limits are per surface, and the surfaces that carry one are listed above.
  • •An API key can record an hourly request rate when it is created. That value is stored with the key but is not applied at the public API edge today, so it is not a control you can rely on — do not use it as a way to throttle a key.
  • •The per-minute figures above can change with deployment sizing. They are what the platform enforces, not a contractual floor.
  • •Self-hosted and on-premise deployments meter differently. Plan allowances are a hosted-platform concept; a self-hosted install is governed by its licence.

ℹ️ Getting the numbers that apply to you

For the limits in force on your own domain, plan, or deployment — and for raising any of them ahead of a launch, a migration, or a load test — contact [email protected]. Include your domain name, the endpoint, and the rate you need. Current plan allowances are readable at any time from the Billing screen in the console, or from GET /api/v1/billing/usage/summary.

Designing for Limits

  • •Send an explicit limit on every list call, and page with a cursor rather than a deep offset.
  • •Prefer one filtered POST /query over pulling a collection and filtering client-side — the filtered call is one request against your allowance instead of many.
  • •Use the aggregate endpoints for counts and sums instead of fetching rows to count them.
  • •Batch writes, but keep batches modest and check the returned counts rather than the status code.
  • •Cache reads that do not change often. Public reads are already cached briefly, so a tight polling loop mostly re-reads the same cached answer while still consuming your allowance.
  • •Handle 429 with backoff and 402 by surfacing it, and never conflate the two in one retry path.
  • •Read your usage on a schedule and alert before you hit a ceiling, rather than discovering it as a failed request on launch day.

Next Steps