PerceptionDocs

Jobs and results

Read a job until it finishes, download the files it produced, and list or delete jobs.

Every document operation creates a job. Its job_id is what the routes below take.

Read a job

Poll the job until status is succeeded, failed or cancelled:

curl https://console.nace.ai/v1/documents/jobs/$JOB_ID \
  -H "Authorization: Bearer $DREX_API_KEY"

The body is the document service's job as it is, with its result once it has succeeded. The result's shape depends on the operation: see Parse, Split, Classify, Extract and Ground.

To skip polling for short jobs, add ?wait_seconds=60 when you create the job. See Wait in the same call.

A job id that isn't yours answers 404, exactly like one that doesn't exist.

Files a job produced

Links in a result to the job's own files point at https://console.nace.ai/v1/documents/jobs/{id}/...: page images, figure crops, transcripts, ground crops, converted PDFs, cell maps, full document content and spreadsheet rows. Fetch them with the same API key.

  • Stored files (page images, transcripts, ground crops, converted PDFs, cell maps) answer 302 to a signed storage link that works for 5 minutes without a key. Let your HTTP client follow the redirect, and don't forward your API key to it: storage refuses a request that carries one, and most clients drop Authorization on a cross-origin redirect anyway.
  • In a browser, add ?redirect=false to get { "url", "expires_at" } instead, since browser code can't read a cross-origin redirect. Then open or download the url. Images, PDFs and JSON open in place; anything else downloads.
  • Everything else (full document content, spreadsheet rows, the request echo and the event stream) is streamed.
  • Rate limits: file downloads don't count against your per-minute limit. The request echo and the event stream do.
  • Expiry: results and their files expire on the document service's retention schedule. After that, a link answers 404 or 410 (code result_expired).

List jobs

GET /v1/documents/jobs lists your jobs, newest first, without their results.

QueryWhat it does
limit1 to 50, default 20.
operationOnly jobs of one operation, such as parse.
statusOnly jobs in one status. It uses the status Drex last recorded, so a job that finished a moment ago may still be listed as running.
cursorThe previous page's next_cursor. Keep going until it's null.

Delete a job

DELETE /v1/documents/jobs/{id} answers 204. A job that's still running is cancelled. The job leaves the list but stays readable by id, and the work it did up to then is still charged.

On this page