Extract
Fill a JSON schema from a document, with a status and citations for every field.
POST /v1/documents/extract fills a JSON Schema from a document. Send the schema inline as schema, or name a saved schema with schema_id, but not both.
curl "https://console.nace.ai/v1/documents/extract?wait_seconds=60" \
-H "Authorization: Bearer $DREX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"source": { "type": "parse_result", "job_id": "'"$PARSE_JOB_ID"'" },
"schema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"total": { "type": "number" }
},
"required": ["invoice_number", "total"]
}
}'Pointing Extract at a finished Parse job, as above, means the document isn't parsed and paid for twice. A url or workspace_file source works too. See Sources.
Options
| Option | Default | What it does |
|---|---|---|
schema | Required, or schema_id | The fields to extract, as a JSON Schema. |
schema_id | Required, or schema | A saved schema's id (sch_...). |
schema_version | Latest | The version of a saved schema to use. Ignored with an inline schema. |
instructions | None | Extra guidance for filling the declared fields. It can't add fields the schema doesn't list. |
page_ranges | All pages | One-based, inclusive ranges. Not available on a parse_result source. |
citations.enabled | true | When true, each field says where in the document it was read. |
citations.include_source_text | true | When true, each citation includes a short quote. |
Result
A finished extract has result.result_type: "extract":
{
"result_type": "extract",
"data": { "invoice_number": "INV-1042", "total": null },
"fields": [
{
"path": "/invoice_number",
"value": "INV-1042",
"status": "found",
"confidence": null,
"citations": [
{ "location": { "kind": "page_region", "page": 1, "polygons": [] }, "source_text": "Invoice INV-1042" }
]
},
{ "path": "/total", "value": null, "status": "not_found", "citations": [] }
],
"schema_id": null,
"schema_version": null,
"units": 1,
"warnings": []
}datais shaped by your schema. A value that wasn't found isnull, never left out.fieldshas one entry per schema field. Join it todataonpath, a JSON Pointer such as/invoice_number.statusisfound,not_foundorambiguous(a value is there but couldn't be chosen uniquely).not_foundis an answer about the document: the job still succeeded.confidencescores the evidence for a found top-level value when there is one. It'snullfor nested paths, arrays and fields that weren't found.citationsgive alocation(see Parse blocks) and an optionalsource_textquote.warningscarry non-fatal notes, such asinstructionsthat no declared field can use.
Saved schemas
Save a schema once, then name it by id instead of sending it with every extract:
curl https://console.nace.ai/v1/extraction-schemas \
-H "Authorization: Bearer $DREX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "invoice",
"description": "Invoice header fields",
"schema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"total": { "type": "number" }
}
}
}'
# 201 {"schema_id":"sch_...","name":"invoice","version":1,"owner":"account","schema":{...},...}Send "schema_id": "sch_..." in place of schema. The extract uses the latest version, or the one you pin with "schema_version": 1, and is billed like any other.
| Route | What it does |
|---|---|
POST /v1/extraction-schemas/{id}/versions | Adds the next version. name and description carry over unless you send them. |
GET /v1/extraction-schemas | Lists the latest version of each schema. |
GET /v1/extraction-schemas/{id} | Reads a schema's latest version. |
GET /v1/extraction-schemas/{id}/versions | Lists a schema's versions. |
GET /v1/extraction-schemas/{id}/versions/{version} | Reads one version. |
- Versions never change. A schema can't be edited or deleted; add a version instead, so an extract that pins
schema_versionalways gets the same schema. - Lists take
limit(1 to 200, default 50) andcursor. - Only your account can see or use your schemas. Another account's
schema_idanswers404 schema_not_found, and the extract isn't billed. - A schema Extract can't use is refused with
422 invalid_schema. - Retries: creating a schema takes no
Idempotency-Key, so a retried create saves a second schema. - Cost: schema routes are free and work at any balance. They count toward the per-minute limit like the other document routes.