Skip to main content
Extend Extract does what Chunkr Extract did: fill a JSON Schema from a document and return citations and confidence for each value. Three things change: how you pass the schema, a few rules the schema must follow, and how you read citations and confidence out of the response.

Create and poll

Citations are opt-in on Extend. Chunkr always returned them. If your application reads citations, set advancedOptions.citationsEnabled: true or they will be missing from the response. This also enables per-field ocrConfidence.
File inputs follow the same rules as Parse: {"url": ...} or {"id": ...} from files.upload. See Migrating Parse.

Schema changes

Chunkr accepted any JSON Schema generated by Pydantic or Zod. Extend validates schemas more strictly. Review your schemas against these rules before migrating:
Pydantic gotcha. Optional[str] generates "anyOf": [{"type": "string"}, {"type": "null"}], which Extend rejects. Either write the schema as a plain dictionary (shown below) or post-process the generated schema to use "type": ["string", "null"]. The Extend TypeScript SDK accepts Zod schemas directly and handles this for you.

Before and after

Extend adds a few typed helpers worth adopting: "extend:type": "date" normalizes dates to yyyy-mm-dd, "extend:type": "currency" returns { amount, iso_4217_currency_code }, and "extend:type": "signature" detects signature blocks. See the schema reference.
Chunkr’s system_prompt parameter maps to config.extractionRules, a plain-language string applied across the whole extraction:

Reading the output

Chunkr returned three parallel objects that mirrored your schema: results, citations, and metrics. Extend returns value (your data) and a flat metadata map keyed by field path.

Field paths

Chunkr’s docs described field paths like line_items[0].description as a way to think about the mirrored structure. On Extend, those paths are literally the keys of metadata.

Confidence

A simple replacement for Chunkr’s Low flag is a threshold on ocrConfidence, for example < 0.8. For a stronger signal, enable the Review Agent and route anything with reviewAgentScore <= 3 to review. See Extend’s confidence guide.

Citations

Chunkr citations carried bboxes[] in {left, top, width, height} pixels. Extend citations carry a polygon[] of {x, y} points. Reduce the polygon to a rectangle and normalize by the page size to get a Chunkr-style box:
Chunkr returned segment-level citations always and word-level when available. On Extend you choose one granularity per run with citationMode. "line" is the default; "word" is closest to Chunkr’s word citations; "block" is closest to segment citations.

Reusing a parsed document

On Chunkr you passed a parse task_id as the file to run several extractions over one parse. On Extend, pass the same file.id to each extract run. Extraction runs parse internally, and parseRunId on the response tells you which parse run was used.
To tune the parse that runs under extraction (the equivalent of Chunkr’s parse_configuration on an extract task), set config.parseConfig with any Parse option.
Production pattern: saved extractors. Instead of sending config on every call, create an Extractor once and reference it by ID. Extractors are versioned, so you can publish a schema change without redeploying code, and you can run evaluation sets against them to measure accuracy on your own documents. See Processors.
Full reference: Extend extraction configuration.

Next Steps

Polling, Webhooks & Production

Async runs, webhooks, retention, and limits.

FAQ

Accounts, deployment, legacy API, and feature differences.