DrexDocs

Limits

Per-request size limits and per-account rate limits, with the exact rules the limiter applies.

Drex applies two kinds of limits. Request limits bound one call to POST /v1/systemone. Rate limits bound how many calls an account can make. Exceeding a request limit returns 422 invalid_request_error. Exceeding a rate limit returns 429 rate_limit_error, or 529 overloaded when a shared pool is full. The messages are listed in Errors.

Request limits

LimitValue
Max questions per request512
Max serialized body (state + questions)1,048,576 bytes
drex-v1.0 state limit (tokens)32,768
drex-v1.0 row limit (state + longest question, tokens)32,768
drex-v1.5 state limit (tokens)131,072
drex-v1.5 row limit (state + longest question, tokens)139,264
Upstream evaluate timeout55s
Route maxDuration60s
Max active API keys3

Max questions per request. The number of keys in questions. Over the limit, the 422 issue has path questions and message at most 512 questions allowed. The minimum is one question.

Max serialized body. Drex serializes state and questions together as JSON and measures the result in UTF-8 bytes. Other top-level fields such as model are not counted. The check runs only after every other validation check has passed, and it is the only issue in the response: path empty, message request body must be at most 1048576 bytes. A 131,072-token state is about 580 KB.

State limit. The state alone, counted in the model's tokenizer the way the model reads it. Over the limit, the 422 issue has path state and message exceeds the 131,072-token state limit, with the model's own number.

Row limit. The state plus the longest question, counted the same way. Other questions don't add to it. Over the limit, the 422 issue has path state and message exceeds the 139,264-token limit, with the model's own number. A body can be under the byte limit and still over a token limit. The example in Errors is a 230 KB state on drex-v1.0.

Upstream evaluate timeout. How long Drex waits for the model service to answer one request. When the wait runs out, the response is 529 overloaded with the message Drex is temporarily unavailable. Retry shortly. and retry-after-ms: 2000. Requests with many questions on a long state take longest, because the model service works through the questions on one GPU.

Route maxDuration. The total time the platform allows one request, including authentication, the rate limiter, and the model call. The evaluate timeout is set below it so the 529 is still delivered.

Max active API keys. Per account. Creating a key at the limit fails on the API Keys page. Revoked keys do not count, so revoke one to make room for another.

Rate limits

TierRPMIn-flight
free1208
paid60016

Both columns apply at the same time. A request must pass the RPM check and the in-flight check to run.

Tiers

An account is free until it has a top-up grant. It is paid from its first completed top-up onward, and stays paid even after that credit is spent or expired. The signup credit, the monthly credit, and a saved card do not make an account paid. Drex caches the tier with the API key check for up to 5 seconds, so the new limits apply within that time after a top-up.

Custom limits

The tier's limits are the default. An account can carry its own limits instead, set by Nace. An account can be unlimited, or can have its own requests per minute, its own in-flight limit, or both. A limit left unset keeps the tier's value. An unlimited account skips the RPM and in-flight checks, and still counts against the shared pools below. Custom limits apply within the same 5 seconds as a tier change.

Requests per minute

The RPM limit is a sliding window, not a calendar minute. Drex records the admission time of every admitted request. A new request is refused when the account already has RPM admissions in the previous 60,000 milliseconds. The refusal is 429 with message Requests per minute exceeded for this account. and retry-after-ms set to the time until the oldest admission leaves the window, never below 1000.

What counts toward RPM on POST /v1/systemone:

  • Every request that passes the API key and credit checks, including requests that then fail body validation, exceed the token limit, or fail at the model service.

What does not count:

  • Requests refused with 401 or 402, which never reach the limiter.
  • Requests refused by any rate limit. A refused request is not recorded anywhere.
  • GET /v1/models, which does not use the limiter.

In flight

A request holds one in-flight slot from admission until Drex has sent its response, success or error, and released the slot. If a slot is never released, for example because the function died, the limiter reclaims it 90 seconds after admission. A request that arrives while the account holds In-flight slots is refused with 429, message Too many concurrent requests for this account., and retry-after-ms: 1000.

Scope

Limits are per account. Every API key on the account draws from the same RPM counter and the same in-flight counter, and so does the dashboard Playground. The Playground runs the same validation, the same limiter call, and the same per-input-token debit as the API. Both validate the body and count its tokens before they call the limiter, so a request that fails validation, or is over the token limit, does not count toward RPM or take an in-flight slot.

Shared capacity

Two pools sit above the per-account limits and are checked after them:

  • An in-flight pool per model fleet, shared by every account. A request counts against the fleet its model and length send it to. Each model has two fleets: a fast one for short requests and a long one for the rest. drex-v1.0 sends a request to its long fleet when the state plus the longest question is over 4,096 tokens. drex-v1.5 sends a request to its long fleet when the state is over 8,192 tokens or the state plus the longest question is over 16,383. Each pool is Drex's promise to its fleet, so a full pool for long requests leaves short ones unaffected.
  • A free-tier in-flight pool shared by every free account, so paid traffic always has room.

When any pool is full, the response is 529 overloaded with message Drex is at capacity. Retry shortly. and retry-after-ms: 2000. The same response is returned when the limiter's store cannot be reached. Pool sizes are operational settings and are not published.

Order of checks

The limiter evaluates account RPM, then account in-flight, then the request's fleet pool, then the free pool, in one atomic step. A request is recorded in all applicable counters only when every check passes. When any check fails, nothing is recorded.

Document routes

The /v1/documents/* routes have a requests-per-minute counter of their own, separate from evaluations, at the account's tier RPM or custom RPM and with the same sliding window. Every request that passes the API key check counts, including reads, lists and deletes; downloads of a job's files under /v1/documents/jobs/{id}/... do not. There is no in-flight limit on these routes. An account with no spendable credit can still read its jobs, at its usual limit. Opening a job's event stream or request echo counts; file downloads don't. Over the limit, the answer is 429 rate_limit_error, code rate_limited, with retry-after. The document service also limits each account and answers the same way. A create body is at most 1,000,000 bytes, and wait_seconds is capped at 60. See Documents.

On this page