Skip to main content
Chunkr Parse and Extend Parse solve the same problem: turn a document into LLM-ready content with layout and position data. The output shape is different enough that you will need to update the code that reads it. This page covers each change.

Create and poll

The Chunkr pattern was create, then poll get until completed. Extend’s SDK wraps both in create_and_poll.
If you need to keep your own polling loop (for example, you store the run ID and poll from a worker), use client.parse_runs.create() and client.parse_runs.retrieve(id). See Polling, Webhooks & Production for details.

File inputs

Chunkr accepted a URL, an uploaded file URL, or a base64 string in a single file parameter. Extend takes an object with either url or id.
Base64 data URIs are not accepted. If you were sending data:application/pdf;base64,..., decode the bytes and use files.upload instead.

Reading the output

Both APIs return output.chunks. Inside a chunk, Chunkr had segments; Extend has blocks.

Field mapping

Segment types to block types

Chunkr classified elements into 18 segment types. Extend uses a smaller set of block types; several Chunkr types collapse into text.
Formula and barcode blocks are opt-in on Extend. Set blockOptions.formulas.enabled and blockOptions.barcodes.readingEnabled in your config. Figures are summarized by a VLM when blockOptions.figures.enabled is set.

Ignoring headers and footers

On Chunkr you set segment_processing={"page_header": {"strategy": "Ignore"}, "page_footer": {"strategy": "Ignore"}}. On Extend, filter blocks client-side when you assemble text:

Bounding boxes

This is the change most likely to break viewer code. The two systems use different box conventions and different units. If you already normalize Chunkr boxes to page fractions (as our docs recommended), the same approach works on Extend. The only change is computing width and height from right and bottom.
Extend auto-rotates skewed pages before parsing, and coordinates are reported in the corrected frame. If you overlay boxes on your original file (not a re-rendered one), check output.metadata.pages[].rotationApplied and follow Extend’s reconciliation guide.

Word-level boxes

Chunkr returned word boxes in pages[].ocr[] by default. On Extend they are opt-in:
Words appear under output.ocr.words[] with content, boundingBox, confidence, pageNumber, and blockId. Because this grows the response, Extend recommends responseType=url for large documents; the output is then delivered as a presigned JSON download in outputUrl.

Page images

Chunkr returned a rendered image per page in pages[].image. Extend does not. If your viewer depends on page images, render them from the original file, which you can download via client.files.retrieve(file_id).presigned_url.

Chunking for RAG

Chunkr chunked by tokens (chunk_processing.target_length, default 4096) and exposed chunk.embed as the embedding text. Extend chunks by characters and exposes chunk.content. The example below uses an explicit 512-token target, a common setting for embedding models.
  • "section" groups by markdown headings and never splits an element across chunks. This is closest to Chunkr’s default behaviour.
  • "page" gives one chunk per page (Extend’s default).
  • "document" returns a single chunk when you do your own splitting downstream.
  • A rough conversion is 4 characters per token, so a 512-token Chunkr target is about maxCharacters: 2000, and Chunkr’s 4096-token default is about maxCharacters: 16000.
Extend has a dedicated Parsing for RAG guide with a ready-to-run config.

Spreadsheets

Chunkr’s ss_* fields exposed cell ranges and formulas on every segment. Extend provides the same provenance through block details when advanced Excel parsing is on.
When targetFormat is html, the same data is also emitted as data-cell and data-formula attributes on table cells. Other useful options: excelSkipHiddenContent, excelUseRawCellValues, and excelSkipCalculation. See Extend’s Excel options.
Chunkr’s configuration lived on the task; Extend’s lives in a config object. Here is how the main options translate.A reasonable starting config that approximates Chunkr’s defaults (set maxCharacters to match your own target_length):
For every option see the Extend parse configuration reference.

Next Steps

Migrating Extract

Schemas, citations, and confidence on Extend.

Polling, Webhooks & Production

Async runs, webhooks, retention, and limits.