Parse
Turn a file into Markdown, plain text, layout blocks and optional chunks.
POST /v1/documents/parse converts a document into structured text. OCR runs when the file needs it, and ocr_applied on the result says whether it did.
curl "https://console.nace.ai/v1/documents/parse?wait_seconds=60" \
-H "Authorization: Bearer $DREX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"source": {
"type": "url",
"url": "https://example.com/report.pdf",
"file_name": "report.pdf"
},
"page_ranges": [{ "start": 1, "end": 3 }],
"output": { "formats": ["markdown", "blocks"] },
"chunking": { "strategy": "page" }
}'The source can be a url, a workspace_file or an earlier parse_result. Reusing a parse skips page_ranges and include_images. See Sources.
Options
| Option | Default | What it does |
|---|---|---|
page_ranges | All pages | One-based, inclusive ranges. Leave it out to parse the whole document. |
parse_mode | low | Effort. low uses native OCR, medium routes to the best parser, high adds page verification and correction. |
output.formats | markdown, blocks | Which renderings to fill on result.document. Formats you leave out are null. text is stripped plain text. |
output.table_format | html | How tables are written inside markdown and text. Blocks keep their own structure. |
output.include_images | false | When true, figure blocks carry a crop link, for PDFs and PNG, JPEG and WebP files. |
output.include_page_markers | true | Inserts --- Page N --- between PDF pages in markdown and text. |
figures.mode | include | Drop figures, keep the crop, or keep it and caption it with describe (one vision call per figure). |
diagrams.mode | omit | Leave flowchart-like figures as images, or rewrite them as Mermaid with mermaid. |
chunking.strategy | none | none leaves chunks empty. page and section produce pieces sized for retrieval. |
spreadsheet | All sheets | Workbook filters for sheet names, hidden cells and formula source text. |
password | None | Password for an encrypted PDF. Never stored or returned. |
Result
A finished parse has result.result_type: "parse". lane names the parser (document, spreadsheet, image, audio or video), and the text is on result.document:
{
"result_type": "parse",
"file_name": "report.pdf",
"lane": "document",
"document": {
"page_count": 96,
"markdown": "# Annual Report\n\n...",
"text": null,
"blocks": [ ... ],
"chunks": []
},
"ocr_applied": false,
"units": 96
}Use markdown to display a document and chunks to index it for retrieval.
Blocks
Each block is a layout element with a location. Its kind depends on the source:
location.kind | Where it points |
|---|---|
page_region | A one-based page and polygons in normalized 0 to 1 coordinates. Use these to draw overlays. |
sheet_range | A workbook tab and an A1 cell range. |
text_span | Inclusive character offsets into the derived text, with page when the span maps to one. |
row_range | A zero-based, inclusive run of rows in a table, with the column when the block is one. |
json_pointer | An RFC 6901 pointer into a JSON source. The empty string is the root. |
Chunks
chunking.strategy: "page" returns one page_region chunk per page. "section" follows the section headings.
Very large results
When a result is too large to store inline, document.markdown and document.text hold a preview, and a content_truncated warning says so:
{ "code": "content_truncated", "target": "document_content", "unit": "characters", "original_size": 120000000, "retained_size": 65000 }Fetch document.content_url for the complete text. For a workbook (target: "spreadsheet_region"), fetch each affected sheet's document.spreadsheet.sheets[].content_url. A first request can answer 202 with status: "pending" and retry-after while the file is prepared. Wait, then send the same request again until it answers 200, and stream that body to a file. Don't save a 202 body as the document.
Large Parquet sheets also carry rows_url, row_count and column_count. GET rows_url?start_row=0&limit=100 reads one page of rows, numbered from zero.
Spreadsheets, audio and video
- Workbooks can include
document.spreadsheet.sheets[].cell_map_url, which maps the rendered table to the sheet's physical rows and columns. Office files converted for layout also link a converted PDF. - Audio and video have
lane: "audio"or"video"and atranscript_url. Reading the job by id also inlines it asresult.transcript: speakers, then orderedspeechandsilencesegments withstart_ms(inclusive) andend_ms(exclusive) that cover the whole recording.
Every one of these links points at https://console.nace.ai/v1/documents/jobs/{id}/... and opens with your API key. See Jobs and results.
Errors
A refused create answers invalid_request_error with the document service's code:
invalid_request: an unknown body field, a bad source, or apasswordon a file that isn't a PDF.unsupported_file_type: Parse doesn't read this type. See Format support.
A file that can't be read, such as a corrupt_file, is accepted but the job ends with status: "failed".
Next, pass the finished job as {"type": "parse_result", "job_id": "..."} to Extract, Split or Ground, so the document isn't parsed and paid for twice.