PerceptionDocs

Extract

Fill a JSON schema from a document, with a status and citations for every field.

POST /v1/documents/extract fills a JSON Schema from a document. Send the schema inline as schema, or name a saved schema with schema_id, but not both.

curl "https://console.nace.ai/v1/documents/extract?wait_seconds=60" \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": { "type": "parse_result", "job_id": "'"$PARSE_JOB_ID"'" },
    "schema": {
      "type": "object",
      "properties": {
        "invoice_number": { "type": "string" },
        "total": { "type": "number" }
      },
      "required": ["invoice_number", "total"]
    }
  }'

Pointing Extract at a finished Parse job, as above, means the document isn't parsed and paid for twice. A url or workspace_file source works too. See Sources.

Options

OptionDefaultWhat it does
schemaRequired, or schema_idThe fields to extract, as a JSON Schema.
schema_idRequired, or schemaA saved schema's id (sch_...).
schema_versionLatestThe version of a saved schema to use. Ignored with an inline schema.
instructionsNoneExtra guidance for filling the declared fields. It can't add fields the schema doesn't list.
page_rangesAll pagesOne-based, inclusive ranges. Not available on a parse_result source.
citations.enabledtrueWhen true, each field says where in the document it was read.
citations.include_source_texttrueWhen true, each citation includes a short quote.

Result

A finished extract has result.result_type: "extract":

{
  "result_type": "extract",
  "data": { "invoice_number": "INV-1042", "total": null },
  "fields": [
    {
      "path": "/invoice_number",
      "value": "INV-1042",
      "status": "found",
      "confidence": null,
      "citations": [
        { "location": { "kind": "page_region", "page": 1, "polygons": [] }, "source_text": "Invoice INV-1042" }
      ]
    },
    { "path": "/total", "value": null, "status": "not_found", "citations": [] }
  ],
  "schema_id": null,
  "schema_version": null,
  "units": 1,
  "warnings": []
}
  • data is shaped by your schema. A value that wasn't found is null, never left out.
  • fields has one entry per schema field. Join it to data on path, a JSON Pointer such as /invoice_number.
  • status is found, not_found or ambiguous (a value is there but couldn't be chosen uniquely). not_found is an answer about the document: the job still succeeded.
  • confidence scores the evidence for a found top-level value when there is one. It's null for nested paths, arrays and fields that weren't found.
  • citations give a location (see Parse blocks) and an optional source_text quote.
  • warnings carry non-fatal notes, such as instructions that no declared field can use.

Saved schemas

Save a schema once, then name it by id instead of sending it with every extract:

curl https://console.nace.ai/v1/extraction-schemas \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "invoice",
    "description": "Invoice header fields",
    "schema": {
      "type": "object",
      "properties": {
        "invoice_number": { "type": "string" },
        "total": { "type": "number" }
      }
    }
  }'
# 201 {"schema_id":"sch_...","name":"invoice","version":1,"owner":"account","schema":{...},...}

Send "schema_id": "sch_..." in place of schema. The extract uses the latest version, or the one you pin with "schema_version": 1, and is billed like any other.

RouteWhat it does
POST /v1/extraction-schemas/{id}/versionsAdds the next version. name and description carry over unless you send them.
GET /v1/extraction-schemasLists the latest version of each schema.
GET /v1/extraction-schemas/{id}Reads a schema's latest version.
GET /v1/extraction-schemas/{id}/versionsLists a schema's versions.
GET /v1/extraction-schemas/{id}/versions/{version}Reads one version.
  • Versions never change. A schema can't be edited or deleted; add a version instead, so an extract that pins schema_version always gets the same schema.
  • Lists take limit (1 to 200, default 50) and cursor.
  • Only your account can see or use your schemas. Another account's schema_id answers 404 schema_not_found, and the extract isn't billed.
  • A schema Extract can't use is refused with 422 invalid_schema.
  • Retries: creating a schema takes no Idempotency-Key, so a retried create saves a second schema.
  • Cost: schema routes are free and work at any balance. They count toward the per-minute limit like the other document routes.

On this page