Create and poll
The Chunkr pattern was create, then pollget until completed. Extend’s SDK wraps both in create_and_poll.
client.parse_runs.create() and client.parse_runs.retrieve(id). See Polling, Webhooks & Production for details.
File inputs
Chunkr accepted a URL, an uploaded file URL, or a base64 string in a singlefile parameter. Extend takes an object with either url or id.
Reading the output
Both APIs returnoutput.chunks. Inside a chunk, Chunkr had segments; Extend has blocks.
Field mapping
Segment types to block types
Chunkr classified elements into 18 segment types. Extend uses a smaller set of block types; several Chunkr types collapse intotext.
Formula and barcode blocks are opt-in on Extend. Set
blockOptions.formulas.enabled and blockOptions.barcodes.readingEnabled in your config. Figures are summarized by a VLM when blockOptions.figures.enabled is set.Ignoring headers and footers
On Chunkr you setsegment_processing={"page_header": {"strategy": "Ignore"}, "page_footer": {"strategy": "Ignore"}}. On Extend, filter blocks client-side when you assemble text:
Bounding boxes
This is the change most likely to break viewer code. The two systems use different box conventions and different units.
If you already normalize Chunkr boxes to page fractions (as our docs recommended), the same approach works on Extend. The only change is computing width and height from
right and bottom.
Word-level boxes
Chunkr returned word boxes inpages[].ocr[] by default. On Extend they are opt-in:
output.ocr.words[] with content, boundingBox, confidence, pageNumber, and blockId. Because this grows the response, Extend recommends responseType=url for large documents; the output is then delivered as a presigned JSON download in outputUrl.
Page images
Chunkr returned a rendered image per page inpages[].image. Extend does not. If your viewer depends on page images, render them from the original file, which you can download via client.files.retrieve(file_id).presigned_url.
Chunking for RAG
Chunkr chunked by tokens (chunk_processing.target_length, default 4096) and exposed chunk.embed as the embedding text. Extend chunks by characters and exposes chunk.content. The example below uses an explicit 512-token target, a common setting for embedding models.
"section"groups by markdown headings and never splits an element across chunks. This is closest to Chunkr’s default behaviour."page"gives one chunk per page (Extend’s default)."document"returns a single chunk when you do your own splitting downstream.- A rough conversion is 4 characters per token, so a 512-token Chunkr target is about
maxCharacters: 2000, and Chunkr’s 4096-token default is aboutmaxCharacters: 16000.
Spreadsheets
Chunkr’sss_* fields exposed cell ranges and formulas on every segment. Extend provides the same provenance through block details when advanced Excel parsing is on.
When
targetFormat is html, the same data is also emitted as data-cell and data-formula attributes on table cells. Other useful options: excelSkipHiddenContent, excelUseRawCellValues, and excelSkipCalculation. See Extend’s Excel options.
Configuration mapping
Configuration mapping
Chunkr’s configuration lived on the task; Extend’s lives in a For every option see the Extend parse configuration reference.
config object. Here is how the main options translate.A reasonable starting config that approximates Chunkr’s defaults (set
maxCharacters to match your own target_length):Next Steps
Migrating Extract
Schemas, citations, and confidence on Extend.
Polling, Webhooks & Production
Async runs, webhooks, retention, and limits.