PerceptionDocs

Ground

Find where each quote appears in a document, with its exact location.

POST /v1/documents/ground finds where each target quote appears in a document. It locates text; it doesn't check whether a claim is true.

curl "https://console.nace.ai/v1/documents/ground?wait_seconds=60" \
  -H "Authorization: Bearer $DREX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "source": {
      "type": "url",
      "url": "https://example.com/report.pdf",
      "file_name": "report.pdf"
    },
    "targets": [
      { "id": "total", "text": "1,200.50" },
      { "id": "signer", "text": "Jane Smith", "hint": "Signed on behalf of the supplier" }
    ]
  }'

The source can be a url, a workspace_file or a parse_result. See Sources.

Options

OptionWhat it does
targets1 to 30 quotes, each with your own id and its text.
targets[].hintNearby text that picks the right occurrence of a quote that appears more than once.
targets[].sheetThe workbook tab to look in.
targets[].semanticfalse (the default) requests the supplied text; true requests a match by meaning. This selects the query type, not a guarantee about the returned match method.
options.max_matchesHow many locations to return per target. Default 10.
options.include_previewsWhen true, each visual hit links a cropped image.

On an HTML source, Ground matches the text a browser would show: script and style contents, hidden elements and attribute values are skipped. On a JSON source, a quote can be a key, a value, or an object pattern that one node must match.

Result

A finished ground has result.result_type: "ground" and one entry in result.targets per request target, in request order:

{
  "result_type": "ground",
  "targets": [
    {
      "id": "total",
      "status": "found",
      "matches": [
        {
          "rank": 1,
          "matched_text": "1,200.50",
          "match_method": "normalized",
          "confidence": null,
          "location": { "kind": "text_span", "page": 1, "char_start": 120, "char_end": 128 },
          "cropped_image_url": null
        }
      ]
    }
  ],
  "units": 1
}

Targets

FieldWhat it holds
idYour id from the request. Join on this, not on the quote text.
statusfound with at least one match, not_found when the quote isn't in the document (the job still succeeded), or failed with an error such as needle_too_short.
matchesLocations, best first. Empty when not_found.

Matches

FieldWhat it holds
rankPosition in matches, starting at 1 for each target. 1 is the place to use. It's a position, not a score.
matched_textThe text actually found, which can differ from the quote when the match isn't exact.
match_methodThe method reported for this hit: exact, normalized or semantic. Visual matching on images and PDFs can report semantic even for a text query. Inspect matched_text and the highlighted region; the query setting alone does not establish an exact match.
confidenceA 0 to 1 locator score when one is available, otherwise null.
locationWhere the quote is: page_region, text_span, sheet_range, json_pointer, jsonl_record, row_range or audio_range. See Parse blocks.
cropped_image_urlA crop of the region with options.include_previews, otherwise null.

For HTML and email, text_span offsets count from the start of the uploaded file, tags and headers included. Crops and transcripts link to https://console.nace.ai/v1/documents/jobs/{id}/... and open with your API key. See Jobs and results.

Ground reads every file type, including audio, video and JSONL. See Format support.

On this page