PerceptionDocs

Parse

Turn a file into Markdown, plain text, layout blocks and optional chunks.

POST /v1/documents/parse converts a document into structured text. OCR runs when the file needs it, and ocr_applied on the result says whether it did.

curl "https://console.nace.ai/v1/documents/parse?wait_seconds=60" \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": {
      "type": "url",
      "url": "https://example.com/report.pdf",
      "file_name": "report.pdf"
    },
    "page_ranges": [{ "start": 1, "end": 3 }],
    "output": { "formats": ["markdown", "blocks"] },
    "chunking": { "strategy": "page" }
  }'

The source can be a url, a workspace_file or an earlier parse_result. Reusing a parse skips page_ranges and include_images. See Sources.

Options

OptionDefaultWhat it does
page_rangesAll pagesOne-based, inclusive ranges. Leave it out to parse the whole document.
parse_modelowEffort. low uses native OCR, medium routes to the best parser, high adds page verification and correction.
output.formatsmarkdown, blocksWhich renderings to fill on result.document. Formats you leave out are null. text is stripped plain text.
output.table_formathtmlHow tables are written inside markdown and text. Blocks keep their own structure.
output.include_imagesfalseWhen true, figure blocks carry a crop link, for PDFs and PNG, JPEG and WebP files.
output.include_page_markerstrueInserts --- Page N --- between PDF pages in markdown and text.
figures.modeincludeDrop figures, keep the crop, or keep it and caption it with describe (one vision call per figure).
diagrams.modeomitLeave flowchart-like figures as images, or rewrite them as Mermaid with mermaid.
chunking.strategynonenone leaves chunks empty. page and section produce pieces sized for retrieval.
spreadsheetAll sheetsWorkbook filters for sheet names, hidden cells and formula source text.
passwordNonePassword for an encrypted PDF. Never stored or returned.

Result

A finished parse has result.result_type: "parse". lane names the parser (document, spreadsheet, image, audio or video), and the text is on result.document:

{
  "result_type": "parse",
  "file_name": "report.pdf",
  "lane": "document",
  "document": {
    "page_count": 96,
    "markdown": "# Annual Report\n\n...",
    "text": null,
    "blocks": [ ... ],
    "chunks": []
  },
  "ocr_applied": false,
  "units": 96
}

Use markdown to display a document and chunks to index it for retrieval.

Blocks

Each block is a layout element with a location. Its kind depends on the source:

location.kindWhere it points
page_regionA one-based page and polygons in normalized 0 to 1 coordinates. Use these to draw overlays.
sheet_rangeA workbook tab and an A1 cell range.
text_spanInclusive character offsets into the derived text, with page when the span maps to one.
row_rangeA zero-based, inclusive run of rows in a table, with the column when the block is one.
json_pointerAn RFC 6901 pointer into a JSON source. The empty string is the root.

Chunks

chunking.strategy: "page" returns one page_region chunk per page. "section" follows the section headings.

Very large results

When a result is too large to store inline, document.markdown and document.text hold a preview, and a content_truncated warning says so:

{ "code": "content_truncated", "target": "document_content", "unit": "characters", "original_size": 120000000, "retained_size": 65000 }

Fetch document.content_url for the complete text. For a workbook (target: "spreadsheet_region"), fetch each affected sheet's document.spreadsheet.sheets[].content_url. A first request can answer 202 with status: "pending" and retry-after while the file is prepared. Wait, then send the same request again until it answers 200, and stream that body to a file. Don't save a 202 body as the document.

Large Parquet sheets also carry rows_url, row_count and column_count. GET rows_url?start_row=0&limit=100 reads one page of rows, numbered from zero.

Spreadsheets, audio and video

  • Workbooks can include document.spreadsheet.sheets[].cell_map_url, which maps the rendered table to the sheet's physical rows and columns. Office files converted for layout also link a converted PDF.
  • Audio and video have lane: "audio" or "video" and a transcript_url. Reading the job by id also inlines it as result.transcript: speakers, then ordered speech and silence segments with start_ms (inclusive) and end_ms (exclusive) that cover the whole recording.

Every one of these links points at https://console.nace.ai/v1/documents/jobs/{id}/... and opens with your API key. See Jobs and results.

Errors

A refused create answers invalid_request_error with the document service's code:

  • invalid_request: an unknown body field, a bad source, or a password on a file that isn't a PDF.
  • unsupported_file_type: Parse doesn't read this type. See Format support.

A file that can't be read, such as a corrupt_file, is accepted but the job ends with status: "failed".

Next, pass the finished job as {"type": "parse_result", "job_id": "..."} to Extract, Split or Ground, so the document isn't parsed and paid for twice.

On this page