> ## Documentation Index
> Fetch the complete documentation index at: https://docs.chunkr.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Migrating Parse

> Move from Chunkr Parse tasks to Extend parse runs

Chunkr Parse and Extend Parse solve the same problem: turn a document into LLM-ready content with layout and position data. The output shape is different enough that you will need to update the code that reads it. This page covers each change.

## Create and poll

The Chunkr pattern was create, then poll `get` until `completed`. Extend's SDK wraps both in `create_and_poll`.

<CodeGroup>
  ```python Chunkr theme={"system"}
  import os
  import time
  from chunkr_ai import Chunkr

  client = Chunkr(api_key=os.environ["CHUNKR_API_KEY"])

  task = client.tasks.parse.create(file="https://example.com/doc.pdf")

  while not task.completed:
      task = client.tasks.parse.get(task_id=task.task_id)
      time.sleep(3)

  if task.status == "Succeeded":
      chunks = task.output.chunks
  ```

  ```python Extend theme={"system"}
  from extend_ai import Extend

  client = Extend()

  run = client.parse_runs.create_and_poll(file={"url": "https://example.com/doc.pdf"})

  if run.status == "PROCESSED":
      chunks = run.output.chunks
  elif run.status == "FAILED":
      print(run.failure_reason, run.failure_message)
  ```

  ```typescript Extend (TypeScript) theme={"system"}
  import { ExtendClient } from "extend-ai";

  const client = new ExtendClient();

  const run = await client.parseRuns.createAndPoll({
    file: { url: "https://example.com/doc.pdf" },
  });

  if (run.status === "PROCESSED") {
    const chunks = run.output?.chunks ?? [];
  } else if (run.status === "FAILED") {
    console.error(run.failureReason, run.failureMessage);
  }
  ```
</CodeGroup>

If you need to keep your own polling loop (for example, you store the run ID and poll from a worker), use `client.parse_runs.create()` and `client.parse_runs.retrieve(id)`. See [Polling, Webhooks & Production](/pages/migrate-to-extend/task-handling) for details.

## File inputs

Chunkr accepted a URL, an uploaded file URL, or a base64 string in a single `file` parameter. Extend takes an object with either `url` or `id`.

<CodeGroup>
  ```python Python theme={"system"}
  # From a URL
  run = client.parse_runs.create_and_poll(file={"url": "https://example.com/doc.pdf"})

  # From a local file (upload first, then reference the ID)
  with open("path/to/doc.pdf", "rb") as f:
      uploaded = client.files.upload(file=f)

  run = client.parse_runs.create_and_poll(file={"id": uploaded.id})
  ```

  ```typescript TypeScript theme={"system"}
  import { createReadStream } from "fs";

  // From a URL
  let run = await client.parseRuns.createAndPoll({
    file: { url: "https://example.com/doc.pdf" },
  });

  // From a local file (upload first, then reference the ID)
  const uploaded = await client.files.upload(createReadStream("path/to/doc.pdf"), {});

  run = await client.parseRuns.createAndPoll({ file: { id: uploaded.id } });
  ```
</CodeGroup>

<Warning>
  Base64 data URIs are not accepted. If you were sending `data:application/pdf;base64,...`, decode the bytes and use `files.upload` instead.
</Warning>

***

## Reading the output

Both APIs return `output.chunks`. Inside a chunk, Chunkr had `segments`; Extend has `blocks`.

<CodeGroup>
  ```json Chunkr theme={"system"}
  {
    "task_id": "...",
    "status": "Succeeded",
    "output": {
      "file_name": "doc.pdf",
      "page_count": 2,
      "chunks": [
        {
          "chunk_id": "...",
          "chunk_length": 312,
          "embed": "## Account Summary\n\n<table>...</table>",
          "segments": [
            {
              "segment_type": "Title",
              "content": "## Account Summary",
              "bbox": { "left": 100, "top": 250, "width": 500, "height": 50 },
              "page_number": 1,
              "page_width": 1224,
              "page_height": 1584
            }
          ]
        }
      ],
      "pages": [ { "page_number": 1, "image": "https://...", "dpi": 144 } ]
    }
  }
  ```

  ```json Extend theme={"system"}
  {
    "object": "parse_run",
    "id": "pr_...",
    "status": "PROCESSED",
    "file": { "id": "file_...", "name": "doc.pdf" },
    "output": {
      "chunks": [
        {
          "object": "chunk",
          "type": "page",
          "content": "# Account Summary\n\n| Date | Description | Amount |...",
          "metadata": { "pageRange": { "start": 1, "end": 1 } },
          "blocks": [
            {
              "object": "block",
              "id": "block_...",
              "type": "heading",
              "content": "# Account Summary",
              "details": {},
              "metadata": { "page": { "number": 1, "width": 612, "height": 792 } },
              "boundingBox": { "left": 56.8, "top": 35.2, "right": 162.2, "bottom": 81.3 },
              "polygon": [ { "x": 56.8, "y": 35.2 }, { "x": 162.2, "y": 35.2 }, ... ]
            }
          ]
        }
      ]
    },
    "metrics": { "pageCount": 2, "processingTimeMs": 8293 },
    "usage": { "credits": 4 }
  }
  ```
</CodeGroup>

### Field mapping

| Chunkr | Extend |
| :- | :- |
| `chunk.embed` | `chunk.content` |
| `chunk.chunk_length` | `len(chunk.content)` (characters) |
| `chunk.segments[]` | `chunk.blocks[]` |
| `segment.segment_type` | `block.type` |
| `segment.content` | `block.content` |
| `segment.page_number` | `block.metadata.page.number` |
| `segment.page_width` / `page_height` | `block.metadata.page.width` / `height` |
| `segment.bbox` | `block.boundingBox` (and `block.polygon`) |
| `segment.image` | `block.details.imageUrl` (figures only) |
| `segment.description` | `block.content` on `figure` blocks |
| `output.page_count` | `metrics.pageCount` |
| `output.file_name` | `file.name` |

### Segment types to block types

Chunkr classified elements into 18 segment types. Extend uses a smaller set of block types; several Chunkr types collapse into `text`.

| Chunkr `segment_type` | Extend `block.type` |
| :- | :- |
| `Title` | `heading` |
| `SectionHeader` | `section_heading` |
| `Text`, `ListItem`, `Caption`, `Footnote`, `Legend`, `LineNumber`, `Unknown` | `text` |
| `Table` | `table` |
| `Picture` | `figure` (with `details.figureType`: `chart`, `image`, `diagram`, `logo`) |
| `FormRegion` | `key_value` |
| `Formula` | `formula` (with `details.latex`) |
| `GraphicalItem` | `barcode` for barcodes and QR codes; `figure` for logos and stamps |
| `PageHeader` | `header` |
| `PageFooter` | `footer` |
| `PageNumber` | `page_number` |
| `Page` | No equivalent |

<Note>
  Formula and barcode blocks are opt-in on Extend. Set `blockOptions.formulas.enabled` and `blockOptions.barcodes.readingEnabled` in your config. Figures are summarized by a VLM when `blockOptions.figures.enabled` is set.
</Note>

### Ignoring headers and footers

On Chunkr you set `segment_processing={"page_header": {"strategy": "Ignore"}, "page_footer": {"strategy": "Ignore"}}`. On Extend, filter blocks client-side when you assemble text:

<CodeGroup>
  ```python Python theme={"system"}
  SKIP = {"header", "footer", "page_number"}

  for chunk in run.output.chunks:
      text = "\n\n".join(b.content for b in chunk.blocks if b.type not in SKIP)
  ```

  ```typescript TypeScript theme={"system"}
  const SKIP = new Set(["header", "footer", "page_number"]);

  for (const chunk of run.output?.chunks ?? []) {
    const text = chunk.blocks
      .filter((b) => !SKIP.has(b.type))
      .map((b) => b.content)
      .join("\n\n");
  }
  ```
</CodeGroup>

***

## Bounding boxes

This is the change most likely to break viewer code. The two systems use different box conventions and different units.

| | Chunkr | Extend |
| :- | :- | :- |
| Shape | `{ left, top, width, height }` | `{ left, top, right, bottom }` plus `polygon[]` |
| Units | Pixels of the rendered page image | Points, in the page's own coordinate space |
| Page size | `segment.page_width`, `segment.page_height` | `block.metadata.page.width`, `block.metadata.page.height` |
| Origin | Top-left | Top-left |

If you already normalize Chunkr boxes to page fractions (as our docs recommended), the same approach works on Extend. The only change is computing width and height from `right` and `bottom`.

<CodeGroup>
  ```python Python theme={"system"}
  def normalize(block):
      """Return a Chunkr-style box normalized to 0-1 page fractions."""
      box = block.bounding_box
      page = block.metadata.page
      return {
          "page_number": page.number,
          "left": box.left / page.width,
          "top": box.top / page.height,
          "width": (box.right - box.left) / page.width,
          "height": (box.bottom - box.top) / page.height,
      }
  ```

  ```typescript TypeScript theme={"system"}
  import { Extend } from "extend-ai";

  function normalize(block: Extend.Block) {
    // Return a Chunkr-style box normalized to 0-1 page fractions
    const box = block.boundingBox;
    const page = block.metadata.page!;
    return {
      pageNumber: page.number,
      left: box.left! / page.width!,
      top: box.top! / page.height!,
      width: (box.right! - box.left!) / page.width!,
      height: (box.bottom! - box.top!) / page.height!,
    };
  }
  ```
</CodeGroup>

<Warning>
  Extend auto-rotates skewed pages before parsing, and coordinates are reported in the corrected frame. If you overlay boxes on your **original** file (not a re-rendered one), check `output.metadata.pages[].rotationApplied` and follow [Extend's reconciliation guide](https://docs.extend.ai/parsing/response-format#reconciling-coordinates-against-your-original-file).
</Warning>

### Word-level boxes

Chunkr returned word boxes in `pages[].ocr[]` by default. On Extend they are opt-in:

```json theme={"system"}
{ "config": { "advancedOptions": { "returnOcr": { "words": true } } } }
```

Words appear under `output.ocr.words[]` with `content`, `boundingBox`, `confidence`, `pageNumber`, and `blockId`. Because this grows the response, Extend recommends `responseType=url` for large documents; the output is then delivered as a presigned JSON download in `outputUrl`.

### Page images

Chunkr returned a rendered image per page in `pages[].image`. Extend does not. If your viewer depends on page images, render them from the original file, which you can download via `client.files.retrieve(file_id).presigned_url`.

***

## Chunking for RAG

Chunkr chunked by **tokens** (`chunk_processing.target_length`, default 4096) and exposed `chunk.embed` as the embedding text. Extend chunks by **characters** and exposes `chunk.content`. The example below uses an explicit 512-token target, a common setting for embedding models.

<CodeGroup>
  ```python Chunkr theme={"system"}
  task = client.tasks.parse.create(
      file=url,
      chunk_processing={"target_length": 512},
  )
  for chunk in task.output.chunks:
      embed(chunk.embed)
  ```

  ```python Extend theme={"system"}
  run = client.parse_runs.create_and_poll(
      file={"url": url},
      config={
          "chunkingStrategy": {
              "type": "section",
              "options": {"minCharacters": 500, "maxCharacters": 2000},
          }
      },
  )
  for chunk in run.output.chunks:
      embed(chunk.content)
  ```
</CodeGroup>

* `"section"` groups by markdown headings and never splits an element across chunks. This is closest to Chunkr's default behaviour.
* `"page"` gives one chunk per page (Extend's default).
* `"document"` returns a single chunk when you do your own splitting downstream.
* A rough conversion is 4 characters per token, so a 512-token Chunkr target is about `maxCharacters: 2000`, and Chunkr's 4096-token default is about `maxCharacters: 16000`.

<Tip>
  Extend has a dedicated [Parsing for RAG](https://docs.extend.ai/parsing/rag) guide with a ready-to-run config.
</Tip>

***

## Spreadsheets

Chunkr's `ss_*` fields exposed cell ranges and formulas on every segment. Extend provides the same provenance through block `details` when advanced Excel parsing is on.

```json theme={"system"}
{
  "config": {
    "advancedOptions": {
      "excelParsingMode": "advanced",
      "excelIncludeCellMetadata": true,
      "excelIncludeCellFormatting": true
    }
  }
}
```

| Chunkr | Extend |
| :- | :- |
| `segment.ss_range` | `block.details.cellReference` (on `table_cell`, `text`, and `heading` blocks) |
| `segment.ss_cells[].formula` | `block.details.formula` |
| `segment.ss_cells[].style` | `block.details.formatting` (`bold`, `italic`, `fontColor`, `backgroundColor`) |
| `segment.ss_header_*` | `blockOptions.tables.tableHeaderContinuationEnabled` |
| `page.ss_sheet_name` | Not exposed |

When `targetFormat` is `html`, the same data is also emitted as `data-cell` and `data-formula` attributes on table cells. Other useful options: `excelSkipHiddenContent`, `excelUseRawCellValues`, and `excelSkipCalculation`. See [Extend's Excel options](https://docs.extend.ai/parsing/configuration#excel).

***

<Accordion title="Configuration mapping">
  Chunkr's configuration lived on the task; Extend's lives in a `config` object. Here is how the main options translate.

  | Chunkr | Extend |
  | :- | :- |
  | Default pipeline | `engine: "parse_performance"` (default). `parse_light` is the faster, cheaper tier; `parse_auto` routes per page. |
  | `pipeline` | Deprecated on Chunkr; drop it. |
  | `chunk_processing.target_length` | `chunkingStrategy.options.maxCharacters` |
  | `chunk_processing.target_length: 0` (one segment per chunk) | Iterate `chunk.blocks` directly |
  | `chunk_processing.tokenizer` | None (character-based) |
  | `segment_processing.table.format: "Markdown"` | `blockOptions.tables.targetFormat: "markdown"` (default is `html`) |
  | `segment_processing.picture.crop_image: "All"` | `blockOptions.figures.figureImageClippingEnabled: true` |
  | `segment_processing.picture.strategy: "LLM"` (default) | `blockOptions.figures.enabled: true` |
  | `segment_processing.picture.strategy: "Auto"` (speed) | `blockOptions.figures.enabled: false` |
  | `segment_processing.text.strategy: "LLM"` (text styling, redlines) | `advancedOptions.formattingDetection: [{ "type": "change_tracking" }]` |
  | `segment_processing.<type>.extended_context` | None. `blockOptions.figures.customInstructions` can steer figure descriptions. |
  | `segment_processing.<type>.strategy: "Ignore"` | Filter `block.type` client-side |
  | `segment_processing.formula` | `blockOptions.formulas.enabled: true` |
  | Charts described as tables inside `Picture` segments | `blockOptions.figures.advancedChartExtractionEnabled: true` |
  | `ocr_strategy: "All"` | OCR always runs. For hard scans and handwriting, `blockOptions.text.agentic.enabled: true`. |
  | `segmentation_strategy: "Page"` | None |
  | `error_handling: "Continue"` | None. Inspect `failureReason` on `FAILED` runs. |
  | `expires_in` | `client.parse_runs.delete(id)` and `client.files.delete(id)` |
  | `file_name` | Taken from the uploaded file |

  A reasonable starting config that approximates Chunkr's defaults (set `maxCharacters` to match your own `target_length`):

  ```json theme={"system"}
  {
    "config": {
      "engine": "parse_performance",
      "chunkingStrategy": { "type": "section", "options": { "maxCharacters": 16000 } },
      "blockOptions": {
        "tables": { "targetFormat": "html" },
        "figures": { "enabled": true, "advancedChartExtractionEnabled": true },
        "formulas": { "enabled": true }
      }
    }
  }
  ```

  For every option see the [Extend parse configuration reference](https://docs.extend.ai/parsing/configuration).
</Accordion>

## Next Steps

<Columns cols={2}>
  <Card title="Migrating Extract" href="/pages/migrate-to-extend/extract" icon="brackets-curly">
    Schemas, citations, and confidence on Extend.
  </Card>

  <Card title="Polling, Webhooks & Production" href="/pages/migrate-to-extend/task-handling" icon="bell">
    Async runs, webhooks, retention, and limits.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.