PerceptionDocs

Classify

Label a document, or each of its pages, with classes you define.

POST /v1/documents/classify reads the original file and scores it against your classes. It doesn't take a parse_result source, so send a url or a workspace_file.

curl "https://console.nace.ai/v1/documents/classify?wait_seconds=60" \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": {
      "type": "url",
      "url": "https://example.com/invoice.pdf",
      "file_name": "invoice.pdf"
    },
    "classes": [
      { "id": "invoice", "label": "Invoice", "description": "A supplier invoice" },
      { "id": "contract", "label": "Contract", "description": "A signed agreement between two parties" }
    ]
  }'

Options

OptionDefaultWhat it does
classesRequiredYour labels. Each needs a stable id, a label and a description the model can use.
granularitydocumentScore the whole file once, or page for one result per page.
page_rangesAll pagesOne-based, inclusive ranges to read and classify.
unknown_policyallowallow may mark a unit unknown. force_best commits an ambiguous unit to one of your classes, though a blank, unreadable or unread page is still unknown.
output.max_alternatives3How many ranked labels to keep per unit, counting the winner.
output.include_reasontrueWhen true, each label may carry a short reason.

Result

result.units has one row for the document, or one per page with granularity: "page":

FieldWhat it holds
unknowntrue when none of your classes won this unit.
pages_read_fractionThe share of expected pages the model read, from 0 to 1.
labelsClasses for this unit, best first. rank: 1 is the assigned class. Later ranks are runners-up.
labels[].class_idOne of your class ids, or the reserved other when none matched part of the unit.
labels[].labelThe class's display name.
labels[].reasonWhy the model chose it, with output.include_reason.
labels[].confidenceThe share of the unit's read pages or sheets given this class, from 0 to 1.

confidence is a share of pages, not a calibrated probability, and rank is a position, not a score. Pages the model never read don't count toward confidence, and the result lists them as classification_unread_pages.

Classify doesn't read audio, video or JSONL files. See Format support.

On this page