Errors and retries
Decide which Drex errors to retry and which to fix, honor retry-after headers, set a long enough timeout, and write a retry loop.
Drex returns one error envelope for every failure, with a type that tells you whether to retry. This guide shows how to classify each error, how long to wait, and how to write the retry loop in the TypeSafe SDK or by hand. For every status, message, and header, see the error reference.
Retry or fix
Every error response has this body. issues appears only on 422, and request_id is always present.
{
"error": {
"type": "rate_limit_error",
"message": "Too many concurrent requests for this account."
},
"request_id": "req_<32 hex characters>"
}Decide by type or by HTTP status.
| HTTP | error.type | What happened | Do this |
|---|---|---|---|
| 429 | rate_limit_error | Your account is over its requests per minute or its in-flight limit | Wait retry-after-ms, then retry |
| 529 | overloaded | Drex is at capacity, the model timed out, or Drex is paused | Wait retry-after-ms, then retry |
| 500 | internal_error | Something failed inside Drex | Retry with backoff |
| No response | Network error or timeout | The request did not complete | Retry with backoff |
| 401 | authentication_error | Missing or invalid nace_sk_ key | Fix the key. See Authentication |
| 402 | insufficient_credit | The account has no spendable credit and no card pays for usage beyond it | Top up or add a card in the dashboard |
| 402 | payment_required | The account has an unpaid invoice | Pay it on the Billing page. Not retryable until it is paid |
| 422 | invalid_request_error | The body failed validation or is over a limit | Fix the body. error.issues says where |
A 4xx that is not 429 will fail the same way on every retry. Retrying it spends your rate limit and delays the real fix.
Honor retry-after
A 429 or 529 carries two headers with the same value: retry-after in whole seconds, rounded up, and retry-after-ms in milliseconds. Read retry-after-ms when you can. Waiting shorter than the header will hit the same limit again. The values Drex sends:
| Cause | error.type | retry-after-ms |
|---|---|---|
| Too many concurrent requests for this account | rate_limit_error | 1000 |
| Requests per minute exceeded for this account | rate_limit_error | Time until the oldest request leaves the 60-second window, at least 1000 |
| Drex is at capacity, or its rate limiter is unavailable | overloaded | 2000 |
| The model timed out or was unavailable | overloaded | 2000 |
| Drex is paused | overloaded | 30000 |
The in-flight limit is the one most clients hit first. A free account may have 8 requests in flight, a paid account 16. If you send a batch of tickets with Promise.all or a thread pool, cap the concurrency at your tier's limit instead of retrying your way through 429s. Limits are on the limits reference.
Set the client timeout to at least 60 seconds
Drex waits up to 55 seconds for the model before it returns a 529, and the route itself stops at 60 seconds. Under load a request can sit in the model's queue for most of that window while capacity is added. A client timeout below 60 seconds abandons requests that were about to succeed, and the retry joins the queue again at the back.
The TypeSafe SDK's default timeout is 10,000 ms per attempt. Raise it.
Configure the TypeSafe SDK
The SDK retries on its own. Its defaults, from RetryPolicy:
maxRetriesis 2, so a request is attempted at most 3 times.- Backoff starts at 500 ms, doubles per attempt, and is capped at 5,000 ms, with up to 25% subtracted as jitter.
- It retries HTTP 408, 429, and every 5xx, which includes 529. It also retries connection errors and timeouts.
- It honors
retry-after-msandretry-afterup to 60,000 ms. Longer server delays fall back to backoff.
Point it at Drex and give it the longer timeout:
import { TypeSafeClient } from "@typesafe-ai/sdk";
export const client = new TypeSafeClient({
apiKey: process.env.DREX_API_KEY,
baseURL: "https://console.nace.ai",
defaultModel: "drex-v1.0",
timeout: 60_000,
retry: { maxRetries: 4, backoffMaxMs: 10_000 },
});Each HTTP status maps to an error class, so you can tell a fix from a retry that ran out:
import {
APIConnectionError,
APIError,
AuthenticationError,
RateLimitError,
UnprocessableEntityError,
noul,
} from "@typesafe-ai/sdk";
import { client } from "./client";
try {
const { answers } = await client.systemOne({
state: "I was charged twice for my March invoice.",
questions: { wants_refund: noul("Is the customer asking for a refund?") },
});
console.log(answers.wants_refund.noul);
} catch (error) {
if (error instanceof UnprocessableEntityError) {
// Fix the request. error.body holds the envelope with error.issues.
console.error("invalid request", error.requestId, error.body);
} else if (error instanceof AuthenticationError) {
console.error("check DREX_API_KEY", error.requestId);
} else if (error instanceof RateLimitError) {
// Retries are used up. error.retryAfterMs is the last server hint.
console.error("rate limited", error.requestId, error.retryAfterMs);
} else if (error instanceof APIError) {
// 402 arrives here as a plain APIError with status 402.
console.error(`Drex ${error.status}`, error.requestId, error.message);
} else if (error instanceof APIConnectionError) {
console.error("network or timeout after retries", error.message);
} else {
throw error;
}
}error.requestId comes from the x-typesafe-request-id header, which Drex sets on every response. See Migrate from TypeSafe for the environment variables that do the same as the constructor options.
Write your own retry loop
Without the SDK, retry 429, 529, 500, and network failures. Use the server's retry-after-ms when present, and capped exponential backoff with jitter otherwise. Stop after a fixed number of attempts, and never retry 401, 402, or 422.
const RETRYABLE = new Set([429, 500, 529]);
export async function systemOne(body: unknown, maxRetries = 4): Promise<unknown> {
for (let attempt = 0; ; attempt++) {
let response: Response | undefined;
try {
response = await fetch("https://console.nace.ai/v1/systemone", {
method: "POST",
headers: {
authorization: `Bearer ${process.env.DREX_API_KEY}`,
"content-type": "application/json",
},
body: JSON.stringify(body),
signal: AbortSignal.timeout(60_000),
});
} catch (error) {
if (attempt >= maxRetries) throw error;
}
if (response) {
if (response.ok) return response.json();
const requestId = response.headers.get("x-request-id");
if (!RETRYABLE.has(response.status) || attempt >= maxRetries) {
throw new Error(`Drex ${response.status} (${requestId}): ${await response.text()}`);
}
console.warn(`Drex ${response.status} (${requestId}), retrying`);
}
await sleep(retryDelayMs(attempt, response?.headers));
}
}
function retryDelayMs(attempt: number, headers?: Headers): number {
const serverMs = Number(headers?.get("retry-after-ms"));
if (serverMs > 0) return serverMs;
const capped = Math.min(500 * 2 ** attempt, 10_000);
return capped / 2 + Math.random() * (capped / 2);
}
const sleep = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms));httpx.TimeoutException is a subclass of httpx.TransportError, so the Python loop retries timeouts too. In both loops the jitter keeps a fleet of clients that hit a 529 at the same moment from retrying at the same moment.
Log and report request_id
Every response, success or failure, carries the same id in three places: the request_id field of the body and the x-request-id and x-typesafe-request-id headers. Log it next to your own correlation id on every call. If a request fails in a way this page does not explain, or a 500 or 529 keeps coming back, write to support@nace.ai with the request_id and the time. Drex joins every log line for a request on that id, so it is the fastest way to a cause.
Know what a retry costs
Drex debits your account only after it has sent a successful response. The debit runs after the response is on the wire and records usage.input_tokens for that request_id. A 401, 402, 422, 429, 500, or 529 is not billed, so retrying a failure does not cost credit. A 429 is decided before Drex reads the body, so a rate-limited request costs nothing but the round trip.
Every successful response is billed, including duplicates. Do not hedge by sending the same request twice in parallel: if both succeed, you pay for both. Usage by request is on the dashboard's Usage page.