PerceptionDocs

Split

Cut a scanned packet or a workbook into the documents it contains, each with a class.

POST /v1/documents/split finds the separate documents inside one file. A PDF becomes page-range segments. A workbook becomes one segment per worksheet, and each sheet is classified against your classes.

curl "https://console.nace.ai/v1/documents/split?wait_seconds=60" \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": {
      "type": "url",
      "url": "https://example.com/packet.pdf",
      "file_name": "packet.pdf"
    },
    "classes": [
      { "id": "invoice", "label": "Invoice", "description": "A supplier invoice" },
      { "id": "receipt", "label": "Receipt", "description": "A payment receipt" }
    ]
  }'

The source can be a url, a workspace_file or a parse_result. See Sources.

Options

OptionDefaultWhat it does
classesRequiredThe document classes to cut on, each with an id, a label and a description.
unknown_policyincludeinclude keeps pages no class matches as unclassified, force gives them the best class, error fails the job.
overlap_policyexclusiveexclusive keeps segments apart. shared_boundary_page puts a boundary page in both neighbouring segments.
split_rulesNoneFree-text guidance on how to cut the packet.
output.include_contentfalseWhen true, each segment carries its Markdown in content.
output.materialize_filesfalseWhen true, each document segment carries a file to download. Workbook segments always come as a single-sheet .xlsx.
page_rangesAll pagesOne-based, inclusive ranges. Not for workbooks.

Result

result.segments is the ordered list of cuts:

FieldWhat it holds
start_page, end_pageThe inclusive page range. For a workbook, both are the sheet's index plus one.
statusmatched when a class won, or unclassified when unknown_policy is include and none did.
classThe winning class id, or null when unclassified.
contentThe segment's Markdown with output.include_content, otherwise null.
artifactsFiles to download when output.materialize_files ran.

To go further, run Extract or Ground on each segment, or on the file it produced.

Split doesn't read audio, video, JSONL or Parquet files. See Format support.

On this page