MCP server
Give a coding agent such as Claude Code, Codex, OpenCode or Cursor Drex's decisions and Perception's document tools, uploads, jobs and saved schemas with nace-mcp, an MCP server.
nace-mcp is an MCP server that gives a coding agent, such as Claude Code, Codex, OpenCode or Cursor, https://console.nace.ai as 21 tools: Drex's decide and list_models, and Perception's five document tools, uploads, jobs, job files and saved extraction schemas. The agent starts it on your machine and talks to it over stdio, and it calls Drex with your API key through the Python SDK.
uvx nace-mcp login # paste a key from the dashboard's API keys page
claude mcp add nace -- uvx nace-mcp # Claude CodeThen ask the agent for what you need, such as "extract the invoice number and total from ~/Downloads/invoice.pdf": it picks the tools and their arguments from the descriptions the server gives it. Install and connect covers Codex, OpenCode, Cursor, Claude Desktop and other ways to install.
Tools
| Tool | What it does | Credits |
|---|---|---|
decide | Ask typed questions about a state and get calibrated probabilities. | Per input token |
list_models | List the models decide can use. | Free |
parse_document | Turn a document into Markdown, text or layout blocks. | 1 per page |
split_document | Cut a packet into page-range segments, each with one of your classes. | 1 per page |
classify_document | Label a document, or each page, with your classes. | 2 per page |
extract_data | Fill a JSON Schema, inline or saved, with a status and a citation for each field. | 3 per page |
ground_items | Find where each quote appears, with its exact location. | 5 per target |
upload_document | Upload a local file into your workspace. | Free |
get_job | Read a job and a summary of its result. | Free |
wait_for_job | Wait for a job to finish. | Free |
list_jobs | List your jobs, newest first. | Free |
cancel_job | Cancel a job, or drop it from the list. | Free |
get_job_request | Read the request a job ran under. | Free |
download_job_file | Save a file a job produced. | Free |
get_job_rows | Read a page of a parsed sheet's rows. | Free |
list_extraction_schemas | List saved extraction schemas. | Free |
get_extraction_schema | Read a saved schema, at its latest version or another. | Free |
list_extraction_schema_versions | List a saved schema's versions. | Free |
create_extraction_schema | Save a new extraction schema. | Free |
add_extraction_schema_version | Add a version to a saved schema. | Free |
get_documentation | Read the usage notes shipped with the server, by topic. | No request |
parse_mode="high" doubles the credits of parse, split, classify and extract. The free tools work at any balance. See Billing.
Install and connect
The commands on this page use uvx, which comes with uv. It runs nace-mcp in an environment of its own, downloading it the first time. To install it once instead, run pipx install nace-mcp (or uv tool install nace-mcp) and use nace-mcp as the command, with no uvx. nace-mcp needs Python 3.11 or later. It checks TLS certificates against your operating system's certificate store, so a proxy certificate your system trusts works too.
| Command | What it does |
|---|---|
nace-mcp | Serve over stdio. Your agent runs it, not you. Stdout carries the protocol, so the server writes nothing else there; its messages go to stderr, which the agent keeps in its MCP log. |
nace-mcp login | Check an API key and save it. See Sign in. |
nace-mcp --version | Print the server and SDK versions. |
nace-mcp -h, --help | Print the usage and exit. |
nace-mcp 0.1.0
nace-sdk 0.1.0Run uvx nace-mcp login once and every agent below finds the key on its own. To pass a key in the agent's config instead, set NACE_API_KEY in the server's environment as each section shows. If an agent started from your desktop can't find uvx, put its full path (from which uvx) in command.
Claude Code and Claude Desktop
claude mcp add nace -- uvx nace-mcpThat adds the server to the current project. claude mcp add --scope user nace -- uvx nace-mcp adds it to every project, and claude mcp add nace --env NACE_API_KEY=nace_sk_... -- uvx nace-mcp passes a key. Run /mcp in Claude Code to see whether it connected.
Claude Desktop reads servers from claude_desktop_config.json (Settings, Developer, Edit Config) when it starts, so restart it after an edit:
{
"mcpServers": {
"nace": {
"command": "uvx",
"args": ["nace-mcp"]
}
}
}Add "env": { "NACE_API_KEY": "nace_sk_..." } beside args to pass a key.
Codex
codex mcp add nace -- uvx nace-mcpAdd --env NACE_API_KEY=nace_sk_... before -- to pass a key. The command writes the server to ~/.codex/config.toml, which the Codex CLI, IDE extension and desktop app share. To edit the file yourself, or to add the server to one trusted project in its .codex/config.toml:
[mcp_servers.nace]
command = "uvx"
args = ["nace-mcp"]
[mcp_servers.nace.env]
NACE_API_KEY = "nace_sk_..."The env table is optional. env_vars = ["NACE_API_KEY"] under [mcp_servers.nace] forwards the variable from the environment Codex runs in instead of writing the key into the file.
OpenCode
OpenCode reads servers from the mcp key of opencode.json, in the project root or at ~/.config/opencode/opencode.json for every project:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"nace": {
"type": "local",
"command": ["uvx", "nace-mcp"],
"enabled": true
}
}
}command holds the command and its arguments as one array. Add "environment": { "NACE_API_KEY": "nace_sk_..." } beside it to pass a key. opencode mcp list shows whether the server connected.
Cursor
Cursor reads servers from .cursor/mcp.json in a project, or ~/.cursor/mcp.json for every project:
{
"mcpServers": {
"nace": {
"type": "stdio",
"command": "uvx",
"args": ["nace-mcp"]
}
}
}Add "env": { "NACE_API_KEY": "${env:NACE_API_KEY}" } beside args to pass the key from your environment, or write the key there in place of ${env:NACE_API_KEY}. The server appears in Cursor's MCP settings, where a toggle turns it off without removing it.
Sign in
nace-mcp login # type the key; it isn't echoed
nace-mcp login --api-key - < key.txt # or read it from stdin| Flag | What it does |
|---|---|
--api-key KEY | The nace_sk_ key. - reads the first line of stdin, or asks at a hidden prompt at a terminal. Omit KEY, or the whole flag, to type the key at a hidden prompt. |
nace-mcp login then:
- Checks that the key looks like a Drex key:
nace_sk_followed by 43 characters. Anything else exits2before any request. - Checks it against Drex with
GET /v1/models, which is free at any balance. It asks$NACE_BASE_URL, else the base URL saved in the config file, elsehttps://console.nace.ai. A refused key exits1with the API's error. So does a host it can't reach, witherror: could not reach <host>: ... Check $NACE_BASE_URL (else base_url in <config path>) and your network; nothing was saved. - Saves it to
~/.nace/config.toml, readable only by you, and printssaved API key to <path>on stderr.
That is the file nace login writes, so signing in with either command signs in both. nace-mcp login has no --base-url: to save another API host, use nace login --base-url, or set NACE_BASE_URL in the server's environment. Without --api-key and without a terminal, nace-mcp login exits 2: pass --api-key KEY or --api-key -. Errors go to stderr as error: ....
Credentials and settings
| Variable | What it does |
|---|---|
NACE_API_KEY | The API key. Wins over api_key in the config file. |
NACE_BASE_URL | The API host. Wins over base_url in the config file; without either, https://console.nace.ai. |
NACE_CONFIG_PATH | Where the config file is, instead of ~/.nace/config.toml. nace-mcp login uses it too. |
NACE_DEFAULT_DECISION_MODEL | The model decide asks when the agent names none. Default drex-latest. |
NACE_LOG_LEVEL | Log each request to stderr, as the CLI does. |
The names from before the packages were renamed from drex-* still work when the new ones are unset or blank: DREX_API_KEY, DREX_BASE_URL, DREX_CONFIG_PATH, DREX_DEFAULT_MODEL and DREX_LOG_LEVEL. So does ~/.drex/config.toml, read when ~/.nace/config.toml doesn't exist and no path variable is set. nace-mcp login always writes the new file.
The agent starts the server, so these come from the agent's environment, not your shell's: set them with env in the JSON above or --env on claude mcp add.
Without a key, the server still starts, and every tool answers with an error: no API key: set $NACE_API_KEY or run `nace-mcp login` . It reads the credentials again on each call, so running nace-mcp login while the agent runs, to sign in or to replace a key Drex refuses, is enough: the next call uses it without a restart. A NACE_API_KEY in the server's environment wins over the saved key and changes only with a restart. An unreadable config file is reported the same way, unless NACE_API_KEY is set.
Ask questions
decide(state, questions, model=None) calls POST /v1/systemone: it asks named, typed questions about one state and returns a calibrated answer to each.
| Argument | What it is |
|---|---|
state | What every question is about: text, or any JSON value such as an object or an array, sent as it is. Text only: images, audio, video and PDFs are refused with 422. |
questions | An object of name to question, up to 512. The name keys the answer, so name it for what it decides (wants_refund, not q1). |
model | The model id. Default $NACE_DEFAULT_DECISION_MODEL, else drex-latest. A pinned id never changes under you; list_models lists them. |
Each question has a type, optional instructions (the question, as text) and criteria:
type | criteria | Answer |
|---|---|---|
noul | Optional: {"true": "...", "false": "..."}, what counts as yes and as no. Needs instructions or a criterion. | noul: the probability of yes, from 0 to 1. |
choice | Required: an object of label to description, or to null. The answer is one of the labels, which have no order. | choice, confidence, and probabilities for each label. |
score | Required: an array of levels, low to high. A level's index is its score. | score (the expected level, which can fall between levels), confidence, legend and probabilities for each level. |
The server checks every question before sending and names the one it refuses, such as question 'topic' (choice): criteria: Field required. The agent passes the arguments as JSON:
{
"state": "I was charged twice for my March invoice.",
"questions": {
"wants_refund": { "type": "noul", "instructions": "Is the customer asking for a refund?" },
"topic": {
"type": "choice",
"instructions": "Which topic is it?",
"criteria": { "billing": "Payments, invoices or charges", "shipping": "Delivery and tracking", "other": null }
},
"urgency": { "type": "score", "instructions": "How urgent is it?", "criteria": ["low", "medium", "high"] }
}
}It gets back the model that answered, one answer per name, the usage and the request id:
{
"model": "drex-v1.5",
"answers": {
"wants_refund": { "type": "noul", "noul": 0.93 },
"topic": {
"type": "choice",
"choice": "billing",
"confidence": 0.88,
"probabilities": { "billing": 0.94, "shipping": 0.01, "other": 0.05 }
},
"urgency": {
"type": "score",
"score": 1.4,
"confidence": 0.5,
"legend": ["low", "medium", "high"],
"probabilities": { "0": 0.1, "1": 0.4, "2": 0.5 }
}
},
"usage": { "input_tokens": 61, "output_tokens": null },
"evaluation_time_ms": 143.0,
"request_id": "req_..."
}model is the pinned version that answered, even when the agent sent an alias. Ask every question about one state in one call: the state is billed once, per input token. The answers are probabilities, not verdicts, so pick a threshold by what each mistake costs. See Questions for what each type accepts and Reading answers for the answers.
List models
list_models() returns {models: [{name, alias_for, release_date, description}]}: the models the key can call, for decide's model. alias_for names the model an alias such as drex-latest serves, and is null for a pinned model. It is free at any balance. See Models.
Sources
The five document tools take the document as exactly one of these:
| Arguments | What it is |
|---|---|
url, with file_name | A public https:// link the document service downloads. See URLs. |
path, with workspace_path and on_conflict | A file on your machine, uploaded into your workspace first. See Local files. |
workspace_id and file_id | A file already in your workspace: what upload_document returns, or a payload's uploaded. |
parse_job_id | One of your finished parse jobs, so the document isn't parsed and paid for again. Not on classify_document, which reads the original file. Not with page_ranges on parse_document or extract_data, parse_mode medium or high, or password. |
Two sources, or none, are refused before any request: pass exactly one source: url, path, workspace_id+file_id, or parse_job_id; got url, path. A text argument sent as an empty string counts as not sent; one with a fixed set of values, such as parse_mode or on_conflict, must be one of them or left out. workspace_id, file_id and parse_job_id must be UUIDs.
URLs
url must start with https://; an http:// link is refused. The document service picks the parser from the file name's extension, so a URL source carries a file_name. The server takes it from the URL's last path segment, without the query and percent-decoded, when that segment has an extension: https://example.com/files/Q3%20report.pdf gives Q3 report.pdf. When the segment has none, the call is refused before any request with can't tell the file type of ... and asks for file_name, such as "report.pdf".
Pass file_name too when the last segment only looks like it has an extension, as in https://arxiv.org/pdf/1706.03762. file_name applies only to a url source.
Local files
path is a file on the machine the server runs on, which is yours: ~ is your home directory, and a relative path is relative to the directory the agent started the server in, so pass an absolute path when in doubt. The file is uploaded first, like upload_document, then read from your workspace. The payload's uploaded holds {workspace_id, file_id, path}: pass that workspace_id and file_id next time instead of uploading again.
| Argument | What it does |
|---|---|
workspace_path | The file's path in your workspace. Default: its file name. Keep the file's extension: the document service types an upload from its first bytes, and from this extension when those don't settle it. |
on_conflict | What happens when a different file already has that path. Left out, or reject, the upload fails with 409 path_conflict and nothing is created. new_version adds a version of that file, which jobs on it that haven't started yet then read. |
Uploading the same bytes to the same path again reuses the file already there. Both arguments apply only to a path source. Every argument the server can check is checked before the upload, so a refused call uploads nothing. When the create fails after the upload went through, the error names the workspace_id and file_id to pass to try again without uploading again. See Same path again.
Arguments every document tool shares
Besides its source, each of the five tools takes:
| Argument | What it does |
|---|---|
parse_mode | How much effort parsing takes: low (the default), medium or high, which doubles the credits. Not on ground_items. medium and high parse again, so not with parse_job_id. |
page_ranges | 1-based pages to read, inclusive, as a string: "3" or "1-5,8". Default: every page. At most 200 ranges and 500 pages. Not on ground_items, and not with parse_job_id on parse_document or extract_data. |
options | Any other request field, as an object. See Request fields. |
idempotency_key | Up to 200 characters. The same key and body return the same job and bill once; the same key with a different body is refused with 409 idempotency_conflict. Default: one minted per call. |
wait_seconds | Hold the create open up to this many seconds, 0 to 60 (default 30), so a short job comes back finished. See Waiting and payloads. |
max_inline_chars | Text or JSON in the result longer than this many characters goes to a file instead (default 20,000). See Spill files. |
A minted key is new on each call: the SDK's retries inside one call return the same job, but calling the tool again without a key creates, and bills, a new one. So when a create fails after a connection failure, a timeout, a 408, a 429 or a 5xx, such as 503 job_pending, the error names the key it carried: To try again, call the tool with idempotency_key="..." and the same arguments. Called again with that key, the tool returns the job if it was created instead of creating, and billing, another. After a path source, the error also names the uploaded file's workspace_id and file_id, to pass in place of path with that key. Pass your own idempotency_key to choose the key. See Retry safely.
Request fields
options goes into the request body as given, for the fields a tool has no argument for: {"name": "March invoices"} for the job's display name, {"project_id": "..."} (a UUID you mint to group jobs), or the tool's own options, such as {"output": {"table_format": "markdown"}}. A key that one of the tool's arguments sets, source included, is refused before any request: options can't set parse_mode: pass it as the tool's own argument instead. On parse_document, formats goes into output.formats, so options may set the other output fields; setting output.formats in options as well as formats is refused. A key named options inside options is a request field: Ground's own options object, as in {"options": {"max_matches": 5}}.
An argument the tool doesn't have is refused before any request, rather than dropped: parse_document has no argument page_range (did you mean page_ranges?), so nothing was sent. Its arguments: ..., followed by a pointer to options for any other request field.
Each tool's options are on its page: Parse, Split, Classify, Extract and Ground. The document service refuses a field it doesn't know, a misspelled one included, with 422.
Waiting and payloads
Each tool creates the job and holds the request open up to wait_seconds. A job that finishes by then comes back with its result in the same call. One still queued or running comes back with "instruction": "the job is still running; call wait_for_job with this job_id", and the agent then calls wait_for_job.
Every tool that returns a job returns the same payload:
{
"job_id": "6f1a2b3c-...",
"kind": "parse",
"status": "succeeded",
"credits": 3,
"usage_final": true,
"result": {
"result_type": "parse",
"file_name": "report.pdf",
"lane": "pdf",
"page_count": 3,
"ocr_applied": false,
"units": 3,
"markdown": { "text": "# Quarterly report\n...", "total_chars": 5120 }
}
}| Field | What it holds |
|---|---|
job_id | The id the job tools take. |
kind | The operation: parse, split, classify, extract or ground. |
status | queued, running, succeeded, failed or cancelled. The last three are final. |
credits, usage_final | The credits the job used, final once usage_final is true. |
error | Why a failed or cancelled job ended: {code, message}. |
result | A summary of the result, made for the agent to read: see each tool below. get_job with full_result=True returns it as the API sent it. |
result_state | expired or not_retained when a finished job's result is gone. |
progress | While the job is queued or running, how far it got, when the document service says. |
instruction | While the job is queued or running, what to call next. |
uploaded | After a path source: {workspace_id, file_id, path}. |
Spill files
Text in a result longer than max_inline_chars, such as a parse's Markdown, comes back as {path, total_chars, head}: the text is written to a file, and head holds its first max_inline_chars characters. Shorter text comes back as {text, total_chars}. JSON longer than that, such as a long list of blocks or segments, comes back as {"spilled": {path, total_chars, head}}, with the JSON text in the file.
The files go in your system's temp directory, named nace-..., and only you can read them. The agent reads or searches them with its own tools. The server never deletes them.
Parse
parse_document turns a document into Markdown, plain text or layout blocks. 1 credit per page.
| Argument | What it does |
|---|---|
formats | A list of markdown, text and blocks, sent as output.formats. Default: the API's markdown and blocks. |
password | An encrypted PDF's password. It is never stored or returned. Not with parse_job_id. |
The result holds file_name, lane, page_count, ocr_applied and units, then:
| Field | What it holds |
|---|---|
markdown, text | Each rendering you asked for, inline or in a spill file. |
content | For a very large document, the full Markdown in a file. See Large documents and workbooks. |
blocks, chunks, ocr | As the API sent them, or spilled when long. |
sheets | A workbook's sheets, each with sheet_name, row_count, column_count, content_url, markdown_deferred, rows_url and cell_map_url when it has them. |
links | Every other job file the result links, such as a transcript_url, for download_job_file. |
disclosures | Notes such as content_truncated. |
Pass the job_id as parse_job_id to split, extract or ground the same document without parsing, and paying for, it again. See Parse.
Large documents and workbooks
For a very large document, the result's markdown and text are only a preview, and document.content_url links the full content. See Very large results. The server then downloads the full Markdown into a new directory only you can read, and returns result.content instead of the preview: {path, bytes, format, head}, where format is markdown. It waits up to 60 seconds for content the document service is still preparing. If the download fails, you get the preview, with preview: true, the content_url, a content_error, and an instruction to save the full content later with download_job_file.
A workbook's sheets each link their own Markdown (content_url) and cell map (cell_map_url): save them with download_job_file. When the sheets carry content_url, because the workbook's rendering is over 1 MiB, markdown and text only preview them, and the server doesn't download them: the result says preview: true, with an instruction to save each sheet with download_job_file(job_id, content_url). A sheet with markdown_deferred: true is only its heading in the preview. A Parquet sheet's rows_url reads with get_job_rows.
Split
split_document cuts a packet of several documents into page-range segments, each with one of your classes. A workbook becomes one segment per sheet. 1 credit per page.
| Argument | What it does |
|---|---|
classes | Required. [{"id", "label", "description"?}], with ids unique and without /. Other class fields, such as instance_key, go in as the API takes them. |
Its options, such as unknown_policy, overlap_policy, split_rules or {"output": {"include_content": true}}, go in options. The result's segments each carry id, sequence, start_page, end_page (and sheet_name for a sheet), status, class_id, class_label, classification_confidence and boundary_confidence, plus instance, content and artifacts when they have them. Segment files, from {"output": {"materialize_files": true}}, save with download_job_file. See Split.
Classify
classify_document labels a document, or each page, with your classes. 2 credits per page.
| Argument | What it does |
|---|---|
classes | Required. [{"id", "label", "description"?, "criteria"?, "subclasses"?}], with ids unique and without /. |
granularity | document (the default) labels the whole file once; page labels each page. |
Classify reads the original file, so it takes no parse_job_id. Options such as unknown_policy or {"output": {"max_alternatives": 3, "include_reason": true}} go in options. The result's units, one per document or page, each carry granularity, page_range, unknown and labels, best first: each with rank (1 is the class assigned), class_id, subclass_id, label, confidence and reason. A label's confidence is the share of pages given that class, not a probability. The result also carries reason_status, unknown_units and pages_read_fraction. A running classify job can already hold a partial result, so wait for a final status. See Classify.
Extract
extract_data fills a JSON Schema from a document, with a status and a citation for each field. 3 credits per page.
| Argument | What it does |
|---|---|
schema | The fields to extract, as an inline JSON Schema. |
schema_id | A saved schema (sch_...) instead, at its latest version. |
schema_version | With schema_id, use this version. |
instructions | Extra guidance for filling the declared fields. It can't add fields. |
Pass exactly one of schema and schema_id; neither, both, or schema_version without schema_id is refused before any request. A finished parse as parse_job_id is the cheapest source. Options such as {"citations": {"enabled": false}} go in options.
The result holds data, shaped by the schema, with null for a value not found. fields lists each field's path, status (found, not_found or ambiguous), confidence, citation_count and first_citation: its location, without polygons, and up to 200 characters of source_text. field_counts counts the fields of each status, and schema_id, schema_version, units and warnings follow when the result has them. not_found is an answer, not an error. For every citation and polygon, read the job with get_job and full_result=True. See Extract.
Ground
ground_items finds where each quote appears in a document, with its exact location. 5 credits per target.
| Argument | What it does |
|---|---|
targets | Required. 1 to 30 of {"id", "text", "hint"?, "sheet"?, "page_hints"?, "semantic"?}. hint is nearby text that picks the right occurrence, sheet a workbook tab, and page_hints the 1-based pages to search first. |
semantic | Deprecated. true sets "semantic": true on every target that doesn't set semantic itself; the request has no switch of its own. On PDF, Word and PowerPoint sources those targets come back failed with semantic_mode_deprecated, while the job still succeeds. Default false. |
Ground locates text; it doesn't check whether a claim is true. It reads the whole document, so it takes neither page_ranges nor parse_mode. Options such as {"options": {"max_matches": 5, "include_previews": true}} go in options.
The result's targets come in request order, each with id, status (found, not_found or failed, with an error), match_count and its first 3 matches. Each match has rank (1 is the place to use), matched_text, match_method (exact, normalized or semantic), confidence, location and cropped_image_url. The result also carries units and transcript_url. Check matched_text and match_method before you trust a match. See Ground.
Uploads
Uploads are free and work at any balance. The bytes go straight to the document service, never through Drex. See Uploads.
Upload a file
upload_document(path, workspace_path=None, on_conflict=None) uploads a local file into your workspace and returns {workspace_id, file_id, path}. Pass the workspace_id and file_id to any tool to work on the file without uploading it again. A tool given a path source uploads the file itself, so the agent needs upload_document only to upload once and then run several tools.
| Argument | What it does |
|---|---|
path | The local file, as in Local files. |
workspace_path | The file's path in your workspace. Default: its file name. Keep its extension. |
on_conflict | What happens when a different file already has that path. Left out, or reject, the upload fails with 409 path_conflict. new_version adds a version of that file, which jobs on it that haven't started yet then read. |
Uploading the same bytes to the same path again returns the file already there. A file of 32 MiB or more goes up in parts.
Jobs
Reading, listing and cancelling jobs is free and works at any balance. See Jobs and results.
Read a job
get_job(job_id, max_inline_chars=20000, full_result=False) reads a job as it is now, without waiting, and returns the payload with a summary of its result. max_inline_chars is as on the tools. full_result=True returns the result exactly as the API sent it, every citation and match included, in a spill file when it is longer than max_inline_chars. A job of another account answers 404, like one that doesn't exist.
Wait for a job
wait_for_job(job_id, timeout=60, max_inline_chars=20000) reads the job every second at first, then less often, up to every 10 seconds, until it finishes or timeout seconds pass. timeout is 0 to 600; anything else is refused. It returns the payload either way:
- A job that finished,
failedandcancelledincluded, comes back with itsresultor itserror. A failed job is not a tool error. - When the time runs out, the job keeps running: the payload has its status,
queuedorrunning, and says"instruction": "the job is still running; call wait_for_job again".
List jobs
list_jobs(operation=None, status=None, cursor=None, limit=None) returns {items, next_cursor}: your jobs, newest first, without their results. Each item has job_id, kind, status, credits, usage_final and created_at.
| Argument | What it does |
|---|---|
operation | Only jobs of one operation: parse, split, classify, extract or ground. |
status | Only jobs in one status: queued, running, succeeded, failed or cancelled. A page can then hold fewer than limit jobs, even none, while next_cursor says more follow. |
limit | Jobs per page, 1 to 50 (default 20). |
cursor | The next_cursor of the page before. Keep going until it is null. |
Cancel a job
cancel_job(job_id) cancels a queued or running job, and drops any job, finished or not, from list_jobs. It returns {job_id, deleted, message}. The job stays readable with get_job, and the work it did is still charged.
Job files
A result links the job's own files: a large document's full Markdown, sheets' Markdown and cell maps, figure crops, ground crops, transcripts, converted PDFs, split segment files and spreadsheet rows. See Files a job produced. Results and their files expire on the document service's schedule; after that, a call on a file fails with 404 or 410 result_expired, while get_job still reads the job, without its result.
Download a file
download_job_file(job_id, path, destination=None, wait_timeout=600) saves a file a job produced and returns {path, bytes}: where it went and its size.
| Argument | What it does |
|---|---|
job_id | The job. |
path | The link from the job's result, as it is, or the part after the job id, such as request. A query or fragment in it is ignored. A path such as .., or a link from another job's result, is refused before any request. |
destination | Where to save the file: a file, which is replaced if it exists, or an existing directory. Default: a new directory in your system's temp directory, named after the job, that only you can read. |
wait_timeout | How long to wait for content the document service is still preparing, in seconds (default 600). |
Without destination, or with a directory, the file is named <job id>.md for a document's full Markdown, else <job id>-<kind>[-<digest>] plus the kind's extension, as nace file names it. Stored files come through their signed link, without your API key; everything else streams from Drex.
Job request
get_job_request(job_id) returns the whole request the job ran under: request_type (the operation), source as the document service stored it, and the options, schema, classes or targets. A PDF password is never returned. Use it to see exactly what a job was asked, for example to run it again with a change. A request longer than 20,000 characters comes back as request_type and a spill file. It expires with the job's result.
Spreadsheet rows
get_job_rows(job_id, path, start_row=0, limit=50) reads one page of a parsed sheet's rows, in file order.
| Argument | What it does |
|---|---|
job_id | The parse job. |
path | The sheet's rows_url from the parse result (only Parquet sources have one), or the part after the job id. |
start_row | The first row, numbered from 0 (default 0). |
limit | Rows per page, 1 to 100 (default 50). |
It returns path, start_row, total_rows, columns (each with name and data_type), rows (each keyed by column name, spilled past 20,000 characters) and next_start_row. Pass next_start_row as start_row for the next page, until it is null.
Saved schemas
Save a JSON Schema once, then name it with schema_id on extract_data instead of sending it every time. Saved schemas are free and work at any balance. See Saved schemas.
List and read schemas
| Tool | What it returns |
|---|---|
list_extraction_schemas(cursor=None, limit=None) | {items, next_cursor}: each schema at its latest version, by schema_id, without its JSON Schema. Items carry schema_id, name, description, version, owner, created_at and updated_at. owner is account for yours and platform for those Nace provides. |
get_extraction_schema(schema_id, version=None) | One schema with its JSON Schema in schema: the latest version, or version. |
list_extraction_schema_versions(schema_id, cursor=None, limit=None) | {items, next_cursor}: the schema's versions, oldest first, each with its schema. |
Lists take limit, 1 to 200 (default 50), and cursor, the next_cursor of the page before; keep going until it is null. A JSON Schema longer than 20,000 characters comes back as a spill file. Another account's schema_id answers 404 schema_not_found.
Save a schema or a version
| Tool | What it does |
|---|---|
create_extraction_schema(name, schema, description=None) | Saves schema, a JSON Schema, as version 1 of a new schema called name, and returns it with its schema_id. |
add_extraction_schema_version(schema_id, schema, name=None, description=None) | Adds the next version to one of your schemas and returns it. Leave out name or description to keep the current one. |
A saved version can't be edited or deleted: add a version instead. An extract with schema_id uses the latest version unless it pins schema_version. Nace's own schemas can't take versions. Neither tool is retried, because a retry could save a second schema or version: after a failure, list the schemas or versions before calling it again.
Usage notes for the agent
get_documentation(topic) returns {topic, text}: usage notes shipped inside the package, so the agent reads what the installed version does rather than what it remembers. Each topic ends with a link to its page in these docs.
topic | What it covers |
|---|---|
overview | What each tool does and costs. The tool's description tells the agent to read it first. |
decide | Question shapes, answers and limits. |
sources | url, path, workspace_id and file_id, parse_job_id, and uploads. |
parse, split, classify, extract, ground | Each tool's arguments, options and summarised result. |
jobs | Waiting, payloads, spill files, job files and rows. |
schemas | Saved extraction schemas. |
errors | What a failed call means and what to do next. |
Any other topic is refused with the list of valid ones.
Retries and timeouts
Each HTTP request may take up to 120 seconds, and a document create wait_seconds more. The server uses the Python SDK's default retry policy, as the CLI does: it retries a 408, a 429, a 5xx (unless the error says "retryable": false), a dropped connection or a timeout up to 2 times. A retry starts only if its wait ends within 30 seconds of the first attempt, so a create that held the request open longer than that and then failed isn't retried.
A document create carries an Idempotency-Key, so these retries never create a second job. Calling the tool again is a new call with a new key, though, unless the agent passes its own: see idempotency_key. After a failure worth retrying, a document tool's error names the key its create carried. create_extraction_schema and add_extraction_schema_version are never retried, because a retry could save a second schema or version.
Handle errors
A failed call comes back as a tool error: text the agent reads, which the MCP library starts with Error executing tool <name>: . For an API error, the text is the status, the error's type and code, its message and the request id. An upload's error comes straight from the document service and has no type, as in 409 path_conflict: .... Up to five fields that failed validation follow, from the error's issues or detail.errors, then (+N more):
Error executing tool parse_document: 422 invalid_request_error/invalid_request: The request does not match the schema for this method. (request req_...)
- output.table_fromat: Extra inputs are not permittedSome errors add a line on what to do next:
| Error | The line added |
|---|---|
401 | Run `nace-mcp login` with a valid key: the server reads the saved key again on each call, so the next call uses it without a restart. Then that a NACE_API_KEY set in the server's environment wins over the saved key, so change it there and restart the server. |
402 insufficient_credit | The account has no spendable credit: top up or add a card at <your Drex host>/dashboard/billing, then call the tool again. |
402 payment_required | The account has an unpaid invoice: pay it at <your Drex host>/dashboard/billing, then call the tool again. A top-up doesn't lift it. |
409 path_conflict | To pass on_conflict="new_version", or another workspace_path to keep both files. |
| A create that failed after a local file was uploaded | The workspace_id and file_id to pass to try again without uploading the file again. |
A document create that failed with a 408, a 429 or a 5xx, or after a connection failure or timeout | To try again, call the tool with idempotency_key="..." and the same arguments: if the job was created, that returns it instead of creating, and billing, another. See idempotency_key. |
On the document tools, these read as follows:
| What happened | What the agent reads |
|---|---|
| A request the document service refused, such as an unsupported file type or an unknown field | 422 invalid_request_error/<code>: ..., then up to five fields from the error's detail.errors. See Errors. |
| A different file at an upload path | 409 path_conflict: ..., without a type. From 32 MiB, Upload did not finish: job <id> is failed (path_conflict: ...). Either way, then the line naming on_conflict="new_version" and workspace_path. |
The same idempotency_key with a different body | 409 conflict_error/idempotency_conflict. |
Content still being prepared after wait_timeout, or content that couldn't be prepared | An error naming the job and path, or 409. |
| A job that failed or was cancelled | Not a tool error: the job, with status and error: {code, message}. |
| A wait that ran out | Not a tool error: the job, still queued or running, with an instruction. |
A failure before any answer from Drex reads differently:
| What happened | What the agent reads |
|---|---|
| No API key | no API key: set $NACE_API_KEY or run `nace-mcp login` . |
| The connection to Drex failed | the connection to Drex failed: ..., then what calling the tool again does. |
| A request timed out | Request timed out (timeout=...). The request may have reached Drex and run., then what calling the tool again does. |
Arguments the server refuses before any request: a decide question it can't send; on the document tools, two sources or none, an http:// URL, a URL without an extension and no file_name, parse_job_id with password, parse_mode medium or high, or page_ranges on parse_document or extract_data, wait_seconds outside 0 to 60, both or neither of schema and schema_id, an options key an argument sets; and on any tool, an argument the tool doesn't have | One line naming the problem, such as question 'topic' (choice): criteria: Field required. An argument the tool doesn't have is named with the argument it was likely meant to be, as in decide has no argument modle (did you mean model?), so nothing was sent. Its arguments: state, questions, model., or with why the tool lacks it: classify_document has no argument parse_job_id (classify reads the original file: ...). Nothing is uploaded or created. |
After a connection failure or a timeout, the request may have reached Drex, so the error says what calling the tool again does:
| Tools | The line added |
|---|---|
Reads, cancel_job and download_job_file | Calling the tool again is safe. |
upload_document, and a document tool whose upload failed | Calling the tool again is safe: the same bytes at the same workspace path return the file already there. |
| A document tool whose create failed | The idempotency_key line above. |
decide | Calling decide again is safe, but if this request reached Drex it was billed, and the new one is billed too. |
create_extraction_schema, add_extraction_schema_version | To check list_extraction_schemas or list_extraction_schema_versions before calling again, or it may be saved twice. |
Quote the request id when you contact support. The error types are in Errors.