Accounts and authentication
| Chunkr | Extend | Notes |
|---|---|---|
| API key from the Chunkr dashboard | API key from the Developers page in the Extend dashboard | Extend keys start with sk_. Chunkr keys do not work on Extend. |
CHUNKR_API_KEY | EXTEND_API_KEY | Extend SDKs read this variable automatically. |
https://api.chunkr.ai | https://api.extend.ai | EU data residency is available at https://api.eu1.extend.ai. |
| No version header | x-extend-api-version: 2026-02-09 | Required on raw HTTP requests. SDKs set it for you. |
pip install chunkr-ai / npm install chunkr-ai | pip install extend-ai / npm install extend-ai | Java and Go SDKs are also available. |
Chunkr(api_key=...) / new Chunkr({ apiKey }) | Extend(token=...) / new ExtendClient({ token }) |
Tasks become runs
| Chunkr | Extend | Notes |
|---|---|---|
| Task | Run | A parse run (pr_...) or extract run (exr_...). |
client.tasks.parse.create(file=...) | client.parse_runs.create_and_poll(file={...}) | Creates the run and polls to completion. Use client.parse_runs.create() to return immediately. |
client.tasks.extract.create(file=..., schema=...) | client.extract_runs.create_and_poll(file={...}, config={"schema": ...}) | |
Poll tasks.parse.get(task_id) until Succeeded | client.parse(...) / client.extract(...) | Extend’s sync endpoints block server-side until done, with a 5-minute cap. Good for testing. Chunkr has no server-side equivalent; polling happens in your code. |
client.tasks.parse.get(task_id) | client.parse_runs.retrieve(id) | |
task.task_id | run.id | |
task.completed | Check run.status against terminal states | PROCESSED, FAILED, CANCELLED. |
client.tasks.delete(task_id) | client.parse_runs.delete(id) | Same for extract_runs. |
client.tasks.cancel(task_id) | client.parse_runs.cancel(id) | Extend can cancel while PROCESSING; Chunkr only cancels tasks still in Starting. |
Status values
| Chunkr | Extend |
|---|---|
Starting | PENDING |
Processing | PROCESSING |
Succeeded | PROCESSED |
Failed | FAILED |
Cancelled | CANCELLED |
File inputs
| Chunkr | Extend | Notes |
|---|---|---|
file="https://..." | file={"url": "https://..."} | |
client.files.create(file=f) then file=uploaded.url | client.files.upload(file=f) then file={"id": uploaded.id} | Extend returns a file_... ID rather than a URL. |
file="data:application/pdf;base64,..." | Not supported | Upload the bytes with files.upload instead. |
file=parse_task.task_id (reuse a parse) | Reuse the same file.id | Extend parses under the hood for extraction; see Migrating Extract. |
file_name parameter | Taken from the upload |
Parse output
| Chunkr | Extend | Notes |
|---|---|---|
output.chunks[] | output.chunks[] | Same top-level idea. |
chunk.segments[] | chunk.blocks[] | |
chunk.embed | chunk.content | Markdown by default. |
chunk.chunk_length (tokens) | len(chunk.content) (characters) | Extend chunk sizes are configured in characters. |
segment.segment_type (Title, Table, …) | block.type (heading, table, …) | Full mapping in Migrating Parse. |
segment.content | block.content | |
segment.bbox (left, top, width, height, pixels) | block.boundingBox (left, top, right, bottom, points) | Also block.polygon for a precise outline. |
segment.page_width / page_height | block.metadata.page.width / height | |
segment.page_number | block.metadata.page.number | |
segment.image (cropped image) | block.details.imageUrl (figures only) | Enable with blockOptions.figures.figureImageClippingEnabled. |
segment.description | block.content on figure blocks | Figure blocks contain the VLM summary. |
pages[].ocr[] (word boxes) | output.ocr.words[] | Enable with advancedOptions.returnOcr.words. |
pages[].image | Not provided | Render pages from the original file; retrieve it with client.files.retrieve(id). |
ss_range, ss_cells | block.details.cellReference, block.details.formula | Enable with advancedOptions.excelParsingMode: "advanced" and excelIncludeCellMetadata. |
output.page_count | metrics.pageCount |
Parse configuration
| Chunkr | Extend | Notes |
|---|---|---|
None (pipeline is deprecated) | config.engine (parse_performance / parse_light / parse_auto) | Extend lets you pick a speed/accuracy tier; Chunkr has one pipeline. |
chunk_processing.target_length | config.chunkingStrategy.options.maxCharacters | Characters, not tokens. |
chunk_processing.tokenizer | None | |
segment_processing.<Type>.format | config.blockOptions.tables.targetFormat | Only tables are configurable; other blocks are markdown. |
segment_processing.<Type>.strategy: Ignore | Filter block.type client-side | |
segment_processing.Picture.crop_image | config.blockOptions.figures.figureImageClippingEnabled | |
segment_processing.<Type>.extended_context | None | Use blockOptions.figures.customInstructions to steer figure descriptions. |
ocr_strategy | Always on | blockOptions.text.agentic.enabled adds VLM OCR correction for hard scans. |
segmentation_strategy: Page | None | |
error_handling: Continue | None | Failed runs report failureReason and failureMessage. |
expires_in | Delete the run and file explicitly | See Data retention. |
Extract output
| Chunkr | Extend | Notes |
|---|---|---|
output.results | output.value | Your schema, populated. |
output.citations (mirrors schema shape) | output.metadata["<path>"].citations[] | Flat map keyed by field path, e.g. line_items[0].total. Opt in with advancedOptions.citationsEnabled. |
output.metrics["<path>"].confidence (High / Low) | output.metadata["<path>"].ocrConfidence (0 to 1) | Also reviewAgentScore (1 to 5) when the Review Agent is enabled. |
citation.bboxes[] (left, top, width, height) | citation.polygon[] ({x, y} points) | Reduce to a rectangle; see Migrating Extract. |
citation.page_number | citation.page.number | |
citation.content | citation.referenceText | |
citation.segment_type / segment_id | None | |
citation.ss_ranges / ss_sheet_name | None in citations | Use Parse with excelIncludeCellMetadata if you need cell provenance. |
system_prompt | config.extractionRules | Plain-language guidance appended to the extraction. |
Webhooks
| Chunkr | Extend | Notes |
|---|---|---|
| Svix-managed endpoints in the Chunkr dashboard | Endpoints under Developers in the Extend dashboard, or via client.webhook_endpoints.create() | |
task.parse.updated (fires on every status change) | parse_run.processed, parse_run.failed | Extend fires on terminal states only. |
| No extract event | extract_run.processed, extract_run.failed | |
| Svix signature headers and libraries | x-extend-request-signature, x-extend-request-timestamp, HMAC-SHA256 | SDK helper: client.webhooks.verify_and_parse(). |
Limits
| Chunkr | Extend | Notes |
|---|---|---|
| 10 files per second | Per-organization limits | See rate limits. |
| 1-hour task timeout | No timeout on async runs; 5 minutes on sync endpoints | |
| 2,000-page soft limit | Use advancedOptions.pageRanges to scope large files | |
SDK retries 429 automatically | Polling helpers back off automatically; add retries for other calls |
Next steps
Getting Started
Create your Extend account and make a first call.
Migrating Parse
Code-level walkthrough of the Parse migration.