DrexDocs

Errors and retries

Decide which Drex errors to retry and which to fix, honor retry-after headers, set a long enough timeout, and write a retry loop.

Drex returns one error envelope for every failure, with a type that tells you whether to retry. This guide shows how to classify each error, how long to wait, and how to write the retry loop in the TypeSafe SDK or by hand. For every status, message, and header, see the error reference.

Retry or fix

Every error response has this body. issues appears only on 422, and request_id is always present.

Error envelope
{
	"error": {
		"type": "rate_limit_error",
		"message": "Too many concurrent requests for this account."
	},
	"request_id": "req_<32 hex characters>"
}

Decide by type or by HTTP status.

HTTPerror.typeWhat happenedDo this
429rate_limit_errorYour account is over its requests per minute or its in-flight limitWait retry-after-ms, then retry
529overloadedDrex is at capacity, the model timed out, or Drex is pausedWait retry-after-ms, then retry
500internal_errorSomething failed inside DrexRetry with backoff
No responseNetwork error or timeoutThe request did not completeRetry with backoff
401authentication_errorMissing or invalid nace_sk_ keyFix the key. See Authentication
402insufficient_creditThe account has no spendable credit and no card pays for usage beyond itTop up or add a card in the dashboard
402payment_requiredThe account has an unpaid invoicePay it on the Billing page. Not retryable until it is paid
422invalid_request_errorThe body failed validation or is over a limitFix the body. error.issues says where

A 4xx that is not 429 will fail the same way on every retry. Retrying it spends your rate limit and delays the real fix.

Honor retry-after

A 429 or 529 carries two headers with the same value: retry-after in whole seconds, rounded up, and retry-after-ms in milliseconds. Read retry-after-ms when you can. Waiting shorter than the header will hit the same limit again. The values Drex sends:

Causeerror.typeretry-after-ms
Too many concurrent requests for this accountrate_limit_error1000
Requests per minute exceeded for this accountrate_limit_errorTime until the oldest request leaves the 60-second window, at least 1000
Drex is at capacity, or its rate limiter is unavailableoverloaded2000
The model timed out or was unavailableoverloaded2000
Drex is pausedoverloaded30000

The in-flight limit is the one most clients hit first. A free account may have 8 requests in flight, a paid account 16. If you send a batch of tickets with Promise.all or a thread pool, cap the concurrency at your tier's limit instead of retrying your way through 429s. Limits are on the limits reference.

Set the client timeout to at least 60 seconds

Drex waits up to 55 seconds for the model before it returns a 529, and the route itself stops at 60 seconds. Under load a request can sit in the model's queue for most of that window while capacity is added. A client timeout below 60 seconds abandons requests that were about to succeed, and the retry joins the queue again at the back.

The TypeSafe SDK's default timeout is 10,000 ms per attempt. Raise it.

Configure the TypeSafe SDK

The SDK retries on its own. Its defaults, from RetryPolicy:

  • maxRetries is 2, so a request is attempted at most 3 times.
  • Backoff starts at 500 ms, doubles per attempt, and is capped at 5,000 ms, with up to 25% subtracted as jitter.
  • It retries HTTP 408, 429, and every 5xx, which includes 529. It also retries connection errors and timeouts.
  • It honors retry-after-ms and retry-after up to 60,000 ms. Longer server delays fall back to backoff.

Point it at Drex and give it the longer timeout:

client.ts
import { TypeSafeClient } from "@typesafe-ai/sdk";

export const client = new TypeSafeClient({
	apiKey: process.env.DREX_API_KEY,
	baseURL: "https://console.nace.ai",
	defaultModel: "drex-v1.0",
	timeout: 60_000,
	retry: { maxRetries: 4, backoffMaxMs: 10_000 },
});

Each HTTP status maps to an error class, so you can tell a fix from a retry that ran out:

handle-errors.ts
import {
	APIConnectionError,
	APIError,
	AuthenticationError,
	RateLimitError,
	UnprocessableEntityError,
	noul,
} from "@typesafe-ai/sdk";
import { client } from "./client";

try {
	const { answers } = await client.systemOne({
		state: "I was charged twice for my March invoice.",
		questions: { wants_refund: noul("Is the customer asking for a refund?") },
	});
	console.log(answers.wants_refund.noul);
} catch (error) {
	if (error instanceof UnprocessableEntityError) {
		// Fix the request. error.body holds the envelope with error.issues.
		console.error("invalid request", error.requestId, error.body);
	} else if (error instanceof AuthenticationError) {
		console.error("check DREX_API_KEY", error.requestId);
	} else if (error instanceof RateLimitError) {
		// Retries are used up. error.retryAfterMs is the last server hint.
		console.error("rate limited", error.requestId, error.retryAfterMs);
	} else if (error instanceof APIError) {
		// 402 arrives here as a plain APIError with status 402.
		console.error(`Drex ${error.status}`, error.requestId, error.message);
	} else if (error instanceof APIConnectionError) {
		console.error("network or timeout after retries", error.message);
	} else {
		throw error;
	}
}

error.requestId comes from the x-typesafe-request-id header, which Drex sets on every response. See Migrate from TypeSafe for the environment variables that do the same as the constructor options.

Write your own retry loop

Without the SDK, retry 429, 529, 500, and network failures. Use the server's retry-after-ms when present, and capped exponential backoff with jitter otherwise. Stop after a fixed number of attempts, and never retry 401, 402, or 422.

drex.ts
const RETRYABLE = new Set([429, 500, 529]);

export async function systemOne(body: unknown, maxRetries = 4): Promise<unknown> {
	for (let attempt = 0; ; attempt++) {
		let response: Response | undefined;
		try {
			response = await fetch("https://console.nace.ai/v1/systemone", {
				method: "POST",
				headers: {
					authorization: `Bearer ${process.env.DREX_API_KEY}`,
					"content-type": "application/json",
				},
				body: JSON.stringify(body),
				signal: AbortSignal.timeout(60_000),
			});
		} catch (error) {
			if (attempt >= maxRetries) throw error;
		}
		if (response) {
			if (response.ok) return response.json();
			const requestId = response.headers.get("x-request-id");
			if (!RETRYABLE.has(response.status) || attempt >= maxRetries) {
				throw new Error(`Drex ${response.status} (${requestId}): ${await response.text()}`);
			}
			console.warn(`Drex ${response.status} (${requestId}), retrying`);
		}
		await sleep(retryDelayMs(attempt, response?.headers));
	}
}

function retryDelayMs(attempt: number, headers?: Headers): number {
	const serverMs = Number(headers?.get("retry-after-ms"));
	if (serverMs > 0) return serverMs;
	const capped = Math.min(500 * 2 ** attempt, 10_000);
	return capped / 2 + Math.random() * (capped / 2);
}

const sleep = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms));

httpx.TimeoutException is a subclass of httpx.TransportError, so the Python loop retries timeouts too. In both loops the jitter keeps a fleet of clients that hit a 529 at the same moment from retrying at the same moment.

Log and report request_id

Every response, success or failure, carries the same id in three places: the request_id field of the body and the x-request-id and x-typesafe-request-id headers. Log it next to your own correlation id on every call. If a request fails in a way this page does not explain, or a 500 or 529 keeps coming back, write to support@nace.ai with the request_id and the time. Drex joins every log line for a request on that id, so it is the fastest way to a cause.

Know what a retry costs

Drex debits your account only after it has sent a successful response. The debit runs after the response is on the wire and records usage.input_tokens for that request_id. A 401, 402, 422, 429, 500, or 529 is not billed, so retrying a failure does not cost credit. A 429 is decided before Drex reads the body, so a rate-limited request costs nothing but the round trip.

Every successful response is billed, including duplicates. Do not hedge by sending the same request twice in parallel: if both succeed, you pay for both. Usage by request is on the dashboard's Usage page.

On this page