Classify
Label a document, or each of its pages, with classes you define.
POST /v1/documents/classify reads the original file and scores it against your classes. It doesn't take a parse_result source, so send a url or a workspace_file.
curl "https://console.nace.ai/v1/documents/classify?wait_seconds=60" \
-H "Authorization: Bearer $DREX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"source": {
"type": "url",
"url": "https://example.com/invoice.pdf",
"file_name": "invoice.pdf"
},
"classes": [
{ "id": "invoice", "label": "Invoice", "description": "A supplier invoice" },
{ "id": "contract", "label": "Contract", "description": "A signed agreement between two parties" }
]
}'Options
| Option | Default | What it does |
|---|---|---|
classes | Required | Your labels. Each needs a stable id, a label and a description the model can use. |
granularity | document | Score the whole file once, or page for one result per page. |
page_ranges | All pages | One-based, inclusive ranges to read and classify. |
unknown_policy | allow | allow may mark a unit unknown. force_best commits an ambiguous unit to one of your classes, though a blank, unreadable or unread page is still unknown. |
output.max_alternatives | 3 | How many ranked labels to keep per unit, counting the winner. |
output.include_reason | true | When true, each label may carry a short reason. |
Result
result.units has one row for the document, or one per page with granularity: "page":
| Field | What it holds |
|---|---|
unknown | true when none of your classes won this unit. |
pages_read_fraction | The share of expected pages the model read, from 0 to 1. |
labels | Classes for this unit, best first. rank: 1 is the assigned class. Later ranks are runners-up. |
labels[].class_id | One of your class ids, or the reserved other when none matched part of the unit. |
labels[].label | The class's display name. |
labels[].reason | Why the model chose it, with output.include_reason. |
labels[].confidence | The share of the unit's read pages or sheets given this class, from 0 to 1. |
confidence is a share of pages, not a calibrated probability, and rank is a position, not a score. Pages the model never read don't count toward confidence, and the result lists them as classification_unread_pages.
Classify doesn't read audio, video or JSONL files. See Format support.