PerceptionDocs

TypeScript SDK

Run Perception's document tools from Node with drex-sdk — parse, split, classify, extract, ground, uploads and saved schemas.

The drex-sdk npm package is the official TypeScript client for https://console.nace.ai. It covers Perception's five tools, uploads, jobs and saved schemas, plus Drex's decision endpoint — see the Drex SDK for systemOne. It needs Node 20 or later and ships as both ESM and CommonJS.

npm install drex-sdk
export DREX_API_KEY="nace_sk_..."

new DrexClient() reads DREX_API_KEY, DREX_BASE_URL and DREX_LOG_LEVEL from the environment. Create a key as shown in Authentication. Each call also takes { signal, timeout, retry, headers } as its last argument.

Parse

client.documents.parse(params, options?)

ParameterTypeRequiredDescription
sourcestring | SourceYesThe document to parse. A plain string must be an https:// URL.
page_rangesArray<{ start: number; end: number }>NoOne-based, inclusive ranges. Default: all pages.
parse_mode"low" | "medium" | "high"NoEffort. low uses native OCR, medium routes to the best parser, high adds page verification and correction. Default: "low".
output{ formats?, table_format?, include_images?, include_page_markers? }Noformats (default ["markdown", "blocks"]), table_format (default "html"), include_images (default false), include_page_markers (default true).
figures{ mode?: "include" | "omit" | "describe" }NoDefault "include"; "describe" captions each figure with one vision call.
diagrams{ mode?: "omit" | "mermaid" }NoDefault "omit"; "mermaid" rewrites flowchart-like figures.
chunking{ strategy?: "none" | "page" | "section" }NoDefault "none"; "page"/"section" produce pieces sized for retrieval.
spreadsheetobjectNoWorkbook filters for sheet names, hidden cells and formula source text. Default: all sheets.
passwordstringNoPassword for an encrypted PDF. Never stored or returned.
wait_secondsnumberNoHold the request open up to 60s to get the result in the same call.
idempotency_keystringNoA retry-safe key; the same key and body return the same job, never billed twice.
parse.ts
import { DrexClient } from "drex-sdk";

const client = new DrexClient();

const job = await client.documents.parse({
  source: "https://example.com/report.pdf",
  page_ranges: [{ start: 1, end: 3 }],
  output: { formats: ["markdown", "blocks"] },
  chunking: { strategy: "page" },
  wait_seconds: 60,
});

console.log(job.result);

See Parse for the result shape.

Split

client.documents.split(params, options?)

ParameterTypeRequiredDescription
sourcestring | SourceYesThe document to split.
classesArray<{ id, label, description? }>YesThe document classes to cut on.
unknown_policy"include" | "force" | "error"No"include" (default) keeps unmatched pages as unclassified, "force" gives them the best class, "error" fails the job.
overlap_policy"exclusive" | "shared_boundary_page"No"exclusive" (default) keeps segments apart; "shared_boundary_page" puts a boundary page in both neighbors.
split_rulesstringNoFree-text guidance on how to cut the packet.
output{ include_content?, materialize_files? }Noinclude_content (default false) puts each segment's Markdown in content; materialize_files (default false) gives each segment a downloadable file.
page_rangesArray<{ start: number; end: number }>NoOne-based, inclusive ranges. Not for workbooks. Default: all pages.
wait_secondsnumberNoHold the request open up to 60s.
idempotency_keystringNoA retry-safe key.
split.ts
const job = await client.documents.split({
  source: "https://example.com/packet.pdf",
  classes: [
    { id: "invoice", label: "Invoice", description: "A supplier invoice" },
    { id: "receipt", label: "Receipt", description: "A payment receipt" },
  ],
  wait_seconds: 60,
});

See Split.

Classify

client.documents.classify(params, options?)

ParameterTypeRequiredDescription
sourcestring | SourceYesThe document to classify. Doesn't take a parse_result source.
classesArray<{ id, label, description? }>YesYour labels.
granularity"document" | "page"NoScore the whole file once, or once per page. Default: "document".
page_rangesArray<{ start: number; end: number }>NoOne-based, inclusive ranges to read and classify. Default: all pages.
unknown_policy"allow" | "force_best"No"allow" (default) may mark a unit unknown; "force_best" commits an ambiguous unit to one of your classes.
output{ max_alternatives?, include_reason? }Nomax_alternatives (default 3) ranked labels per unit; include_reason (default true) adds a short reason per label.
wait_secondsnumberNoHold the request open up to 60s.
idempotency_keystringNoA retry-safe key.
classify.ts
const job = await client.documents.classify({
  source: "https://example.com/invoice.pdf",
  classes: [
    { id: "invoice", label: "Invoice", description: "A supplier invoice" },
    { id: "contract", label: "Contract", description: "A signed agreement between two parties" },
  ],
  granularity: "page",
  wait_seconds: 60,
});

See Classify.

Extract

client.documents.extract(params, options?)

ParameterTypeRequiredDescription
sourcestring | SourceYesThe document to extract from. Reuse a finished parse with a parse_result source so it isn't parsed and paid for twice.
schemaRecord<string, unknown>YesThe fields to extract, as a JSON Schema. A saved schema's schema_id is accepted by the server but not yet by this field — see below.
instructionsstringNoExtra guidance for filling the declared fields. Can't add fields the schema doesn't list.
page_rangesArray<{ start: number; end: number }>NoOne-based, inclusive ranges. Not available on a parse_result source. Default: all pages.
citations{ enabled?, include_source_text? }Noenabled (default true) says where in the document each field was read; include_source_text (default true) adds a short quote.
wait_secondsnumberNoHold the request open up to 60s.
idempotency_keystringNoA retry-safe key.
extract.ts
const parsed = await client.documents.parse({
  source: "https://example.com/invoice.pdf",
  wait_seconds: 60,
});

const job = await client.documents.extract({
  source: { type: "parse_result", job_id: parsed.job_id },
  schema: {
    type: "object",
    properties: {
      invoice_number: { type: "string" },
      total: { type: "number" },
    },
    required: ["invoice_number", "total"],
  },
  wait_seconds: 60,
});

console.log((job.result as { data: unknown }).data); // { invoice_number: 'INV-1042', total: null }

See Extract for the result's fields, status and citations.

Ground

client.documents.ground(params, options?)

ParameterTypeRequiredDescription
sourcestring | SourceYesThe document to search.
targetsArray<{ id, text, hint?, sheet? }>Yes1 to 30 quotes. hint (nearby text, for a repeated quote) and sheet (workbook tab) are optional per target.
semanticbooleanNoWhen true, also finds a quote by meaning, not only as a literal string. Default: false.
options{ max_matches?, include_previews? }Nomax_matches (default 10) locations per target; include_previews (default false) links a cropped image per visual hit.
wait_secondsnumberNoHold the request open up to 60s.
idempotency_keystringNoA retry-safe key.
ground.ts
const job = await client.documents.ground({
  source: "https://example.com/report.pdf",
  targets: [
    { id: "total", text: "1,200.50" },
    { id: "signer", text: "Jane Smith", hint: "Signed on behalf of the supplier" },
  ],
  wait_seconds: 60,
});

See Ground.

Upload a local file

client.documents.upload(content, params, options?)

ParameterTypeRequiredDescription
contentBlob | Uint8Array | ArrayBufferYesThe file's bytes.
params.file_namestringYesThe file's name, sent with the upload.
params.pathstringNoPath in the workspace. Default: file_name.
upload.ts
import { readFile } from "node:fs/promises";
import { workspaceFile } from "drex-sdk";

const file = await client.documents.upload(await readFile("invoice.pdf"), {
  file_name: "invoice.pdf",
  path: "invoices/invoice.pdf",
});

const job = await client.documents.extract({
  source: workspaceFile(file),
  schema: { type: "object", properties: { total: { type: "number" } } },
  wait_seconds: 60,
});

The file goes straight to the document service with a one-use token; your API key is never sent with it. Files of 32 MiB or more are uploaded in parts automatically. See Uploads.

List, read and delete jobs

MethodParametersDescription
client.jobs.get(jobId)jobId: stringOne job and, once it's succeeded, its result.
client.jobs.list(params)limit?, operation?, status?, cursor?One page of jobs, newest first, without results.
client.jobs.iter(params?)operation?, status?Async-iterate every job across every page.
client.jobs.events(jobId, opts?)jobId: string, opts.lastEventId?: stringAsync-iterate progress as Server-Sent Events, for up to five minutes.
client.jobs.fileLink(jobId, path)jobId: string, path: stringA five-minute signed link to a file the job produced.
client.jobs.delete(jobId)jobId: stringCancels a running job and drops it from the list; billed work stays charged.
await client.jobs.list({ limit: 20, operation: "parse", status: "succeeded" });
for await (const job of client.jobs.iter()) { /* ... */ }
for await (const event of client.jobs.events(jobId)) console.log(event.status);
const { url } = await client.jobs.fileLink(jobId, "parse-images/<ref>");
await client.jobs.delete(jobId);

See Jobs and results.

Save a reusable extraction schema

Save a JSON Schema once, then reuse it across extracts without sending it inline. Saved schemas can't be edited; add a version instead.

MethodParametersDescription
client.extractionSchemas.create(params)name: string, schema: object, description?: stringSaves version 1 of a new schema. Not idempotent: a retried create saves a second schema.
client.extractionSchemas.list(params?)limit?, cursor?One page of the account's saved schemas, at their latest version.
client.extractionSchemas.get(schemaId)schemaId: stringThe schema's latest version.
client.extractionSchemas.createVersion(schemaId, params)schemaId: string, schema: object, name?, description?Adds a new version. Omit name/description to keep the current ones.
client.extractionSchemas.listVersions(schemaId, params?)schemaId: string, limit?, cursor?One page of the schema's versions, newest first.
client.extractionSchemas.getVersion(schemaId, version)schemaId: string, version: numberOne specific version.
const saved = await client.extractionSchemas.create({
  name: "invoice",
  description: "Invoice header fields",
  schema: {
    type: "object",
    properties: { invoice_number: { type: "string" }, total: { type: "number" } },
  },
});
await client.extractionSchemas.list({ limit: 20 });
await client.extractionSchemas.get(saved.schema_id);
await client.extractionSchemas.createVersion(saved.schema_id, {
  schema: {
    type: "object",
    properties: { invoice_number: { type: "string" }, total: { type: "number" }, due_date: { type: "string" } },
  },
});
await client.extractionSchemas.listVersions(saved.schema_id);

extract()'s schema field is required and doesn't yet accept a saved schema's schema_id — that field is server-side only for now. See Extract's saved schemas.

Handle errors

Every error throws a DrexError; HTTP failures throw an APIError subclass with status, body, requestId, type, code and serverRetryable — see Handle errors for the full class table and retry behavior. The document routes add their own codes, such as 503 document_billing_not_configured and 422 invalid_source, and their own per-minute rate limit separate from systemOne's — see Errors.

On this page