Python SDK
Run Perception's document tools from Python with drex-sdk — parse, split, classify, extract, ground, uploads and saved schemas.
drex-sdk is the official Python client for https://console.nace.ai. It covers Perception's five tools, uploads, jobs and saved schemas, plus Drex's decision endpoint — see the Drex SDK for system_one. It needs Python 3.11 or later and depends on httpx and pydantic.
pip install drex-sdk
export DREX_API_KEY="nace_sk_..."DrexClient() reads DREX_API_KEY, DREX_BASE_URL and DREX_LOG_LEVEL from the environment. Create a key as shown in Authentication. Every method below also takes retry, timeout and extra_headers, and every create also takes extra_body, for one call.
Parse
client.documents.parse(source, **options)
| Parameter | Type | Required | Description |
|---|---|---|---|
source | str | UrlSource | WorkspaceFileSource | ParseResultSource | Yes | The document to parse. A plain str must be an https:// URL. |
page_ranges | list[dict] | No | One-based, inclusive ranges, e.g. [{"start": 1, "end": 3}]. Default: all pages. |
parse_mode | "low" / "medium" / "high" | No | Effort. low uses native OCR, medium routes to the best parser, high adds page verification and correction. Default: low. |
output | dict | No | formats (default ["markdown", "blocks"]; add "text"), table_format (default "html"), include_images (default False), include_page_markers (default True). |
figures | dict | No | mode: "include" (default), "omit", or "describe" to caption each figure with one vision call. |
diagrams | dict | No | mode: "omit" (default) or "mermaid" to rewrite flowchart-like figures. |
chunking | dict | No | strategy: "none" (default), "page" or "section", sized for retrieval. |
spreadsheet | dict | No | Workbook filters for sheet names, hidden cells and formula source text. Default: all sheets. |
password | str | No | Password for an encrypted PDF. Never stored or returned. |
wait_seconds | int | No | Hold the request open up to 60s to get the result in the same call. Default: None (don't wait). |
idempotency_key | str | No | A retry-safe key; the same key and body return the same job, never billed twice. |
from drex_sdk import DrexClient
with DrexClient() as client:
job = client.documents.parse(
"https://example.com/report.pdf",
page_ranges=[{"start": 1, "end": 3}],
output={"formats": ["markdown", "blocks"]},
chunking={"strategy": "page"},
wait_seconds=60,
)
print(job.result["document"]["markdown"])See Parse for the result shape.
Split
client.documents.split(source, *, classes, **options)
| Parameter | Type | Required | Description |
|---|---|---|---|
source | str | UrlSource | WorkspaceFileSource | ParseResultSource | Yes | The document to split. |
classes | list[dict] | Yes | The document classes to cut on, each {"id", "label", "description"}. |
unknown_policy | "include" / "force" / "error" | No | include (default) keeps unmatched pages as unclassified, force gives them the best class, error fails the job. |
overlap_policy | "exclusive" / "shared_boundary_page" | No | exclusive (default) keeps segments apart; shared_boundary_page puts a boundary page in both neighbors. |
split_rules | str | No | Free-text guidance on how to cut the packet. |
output | dict | No | include_content (default False) puts each segment's Markdown in content; materialize_files (default False) gives each segment a downloadable file. |
page_ranges | list[dict] | No | One-based, inclusive ranges. Not for workbooks. Default: all pages. |
wait_seconds | int | No | Hold the request open up to 60s. Default: None. |
idempotency_key | str | No | A retry-safe key. |
job = client.documents.split(
"https://example.com/packet.pdf",
classes=[
{"id": "invoice", "label": "Invoice", "description": "A supplier invoice"},
{"id": "receipt", "label": "Receipt", "description": "A payment receipt"},
],
wait_seconds=60,
)
for segment in job.result["segments"]:
print(segment["start_page"], segment["end_page"], segment["class"])See Split.
Classify
client.documents.classify(source, *, classes, **options)
| Parameter | Type | Required | Description |
|---|---|---|---|
source | str | UrlSource | WorkspaceFileSource | Yes | The document to classify. Doesn't take a parse_result source. |
classes | list[dict] | Yes | Your labels, each {"id", "label", "description"}. |
granularity | "document" / "page" | No | Score the whole file once, or once per page. Default: "document". |
page_ranges | list[dict] | No | One-based, inclusive ranges to read and classify. Default: all pages. |
unknown_policy | "allow" / "force_best" | No | allow (default) may mark a unit unknown; force_best commits an ambiguous unit to one of your classes. |
output | dict | No | max_alternatives (default 3) ranked labels per unit; include_reason (default True) adds a short reason per label. |
wait_seconds | int | No | Hold the request open up to 60s. Default: None. |
idempotency_key | str | No | A retry-safe key. |
job = client.documents.classify(
"https://example.com/invoice.pdf",
classes=[
{"id": "invoice", "label": "Invoice", "description": "A supplier invoice"},
{"id": "contract", "label": "Contract", "description": "A signed agreement between two parties"},
],
granularity="page",
wait_seconds=60,
)See Classify.
Extract
client.documents.extract(source, *, schema, instructions=None, **options)
| Parameter | Type | Required | Description |
|---|---|---|---|
source | str | UrlSource | WorkspaceFileSource | ParseResultSource | Yes | The document to extract from. Reuse a finished parse with ParseResultSource so it isn't parsed and paid for twice. |
schema | dict | Yes | The fields to extract, as a JSON Schema. A saved schema's schema_id is accepted by the server but not yet by this parameter — see below. |
instructions | str | No | Extra guidance for filling the declared fields. Can't add fields the schema doesn't list. |
page_ranges | list[dict] | No | One-based, inclusive ranges. Not available on a parse_result source. Default: all pages. |
citations | dict | No | enabled (default True) says where in the document each field was read; include_source_text (default True) adds a short quote. |
wait_seconds | int | No | Hold the request open up to 60s. Default: None. |
idempotency_key | str | No | A retry-safe key. |
from drex_sdk import DrexClient, ParseResultSource
parsed = client.documents.parse("https://example.com/invoice.pdf", wait_seconds=60)
job = client.documents.extract(
ParseResultSource(job_id=parsed.job_id),
schema={
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"total": {"type": "number"},
},
"required": ["invoice_number", "total"],
},
wait_seconds=60,
)
print(job.result["data"]) # {'invoice_number': 'INV-1042', 'total': None}See Extract for the result's fields, status and citations.
Ground
client.documents.ground(source, *, targets, **options)
| Parameter | Type | Required | Description |
|---|---|---|---|
source | str | UrlSource | WorkspaceFileSource | ParseResultSource | Yes | The document to search. |
targets | list[dict] | Yes | 1 to 30 quotes, each {"id", "text"}. hint (nearby text, for a repeated quote) and sheet (workbook tab) are optional per target. |
semantic | bool | No | When True, also finds a quote by meaning, not only as a literal string. Default: False. |
options | dict | No | max_matches (default 10) locations per target; include_previews (default False) links a cropped image per visual hit. |
wait_seconds | int | No | Hold the request open up to 60s. Default: None. |
idempotency_key | str | No | A retry-safe key. |
job = client.documents.ground(
"https://example.com/report.pdf",
targets=[
{"id": "total", "text": "1,200.50"},
{"id": "signer", "text": "Jane Smith", "hint": "Signed on behalf of the supplier"},
],
wait_seconds=60,
)
for target in job.result["targets"]:
print(target["id"], target["status"], target["matches"][:1])See Ground.
Upload a local file
client.documents.upload(content, *, path=None, file_name=None)
| Parameter | Type | Required | Description |
|---|---|---|---|
content | str | Path | bytes | BinaryIO | Yes | The file, or a path to it. |
path | str | No | Path in the workspace. Default: the file's own name. |
file_name | str | No | Overrides the name read from content, e.g. when content is raw bytes. |
file = client.documents.upload("invoice.pdf", path="invoices/invoice.pdf")
job = client.documents.extract(file.as_source(), schema={"type": "object", "properties": {"total": {"type": "number"}}}, wait_seconds=60)The file goes from your machine straight to the document service with a one-use token; your API key is never sent with it. Files of 32 MiB or more are uploaded in parts automatically. See Uploads.
List, read and delete jobs
| Method | Parameters | Description |
|---|---|---|
client.jobs.get(job_id) | job_id: str | One job and, once it's succeeded, its result. |
client.jobs.list(**options) | limit: int, operation: str, status: str, cursor: str (all optional) | One page of jobs, newest first, without results. |
client.jobs.iter(**options) | operation: str, status: str (optional) | Every job across every page. |
client.jobs.wait(job_id, **options) | job_id: str, timeout: float (default 60), raise_on_failure: bool (default True) | Poll until the job is terminal. |
client.jobs.events(job_id) | job_id: str, last_event_id: str (optional) | Stream progress as Server-Sent Events, for up to five minutes. |
client.jobs.file_link(job_id, path) | job_id: str, path: str | A five-minute signed link to a file the job produced. |
client.jobs.download(job_id, path, destination) | job_id: str, path: str, destination: str | Path | Saves a file the job produced to destination. |
client.jobs.delete(job_id) | job_id: str | Cancels a running job and drops it from the list; billed work stays charged. |
client.jobs.list(limit=20, operation="parse", status="succeeded")
for job in client.jobs.iter():
...
for event in client.jobs.events(job_id):
print(event.status)
client.jobs.download(job_id, "parse-images/<ref>", "page-1.png")
client.jobs.delete(job_id)See Jobs and results.
Save a reusable extraction schema
Save a JSON Schema once, then reuse it across extracts without sending it inline. Saved schemas can't be edited; add a version instead.
| Method | Parameters | Description |
|---|---|---|
client.extraction_schemas.create(name, schema, **options) | name: str, schema: dict, description: str (optional) | Saves version 1 of a new schema. Not idempotent: a retried create saves a second schema. |
client.extraction_schemas.list(**options) | limit: int, cursor: str (optional) | One page of the account's saved schemas, at their latest version. |
client.extraction_schemas.get(schema_id) | schema_id: str | The schema's latest version. |
client.extraction_schemas.create_version(schema_id, schema, **options) | schema_id: str, schema: dict, name: str, description: str (optional) | Adds a new version. Omit name/description to keep the current ones. |
client.extraction_schemas.list_versions(schema_id, **options) | schema_id: str, limit: int, cursor: str (optional) | One page of the schema's versions, newest first. |
client.extraction_schemas.get_version(schema_id, version) | schema_id: str, version: int | One specific version. |
saved = client.extraction_schemas.create(
"invoice",
{
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"total": {"type": "number"},
},
},
description="Invoice header fields",
)
client.extraction_schemas.list(limit=20)
client.extraction_schemas.get(saved.schema_id)
client.extraction_schemas.create_version(
saved.schema_id,
{"type": "object", "properties": {"invoice_number": {"type": "string"}, "total": {"type": "number"}, "due_date": {"type": "string"}}},
)
client.extraction_schemas.list_versions(saved.schema_id)documents.extract()'s schema parameter is required and doesn't yet accept a saved schema's schema_id — that field is server-side only for now. See Extract's saved schemas.
Handle errors
Every error raises a DrexError; HTTP failures raise a DrexAPIError subclass with status, body, request_id, type, code and server_retryable — see Handle errors for the full class table and retry behavior. The document routes add their own codes, such as 503 document_billing_not_configured and 422 invalid_source, and their own per-minute rate limit separate from system_one's — see Errors.