> ## Documentation Index
> Fetch the complete documentation index at: https://docs.chunkr.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Migrating Extract

> Move from Chunkr Extract tasks to Extend extract runs

Extend Extract does what Chunkr Extract did: fill a JSON Schema from a document and return citations and confidence for each value. Three things change: how you pass the schema, a few rules the schema must follow, and how you read citations and confidence out of the response.

## Create and poll

<CodeGroup>
  ```python Chunkr theme={"system"}
  import os
  import time
  from chunkr_ai import Chunkr

  client = Chunkr(api_key=os.environ["CHUNKR_API_KEY"])

  task = client.tasks.extract.create(
      file="https://example.com/invoice.pdf",
      schema=schema,
  )

  while not task.completed:
      task = client.tasks.extract.get(task_id=task.task_id)
      time.sleep(3)

  if task.status == "Succeeded":
      data = task.output.results
  ```

  ```python Extend theme={"system"}
  from extend_ai import Extend

  client = Extend()

  run = client.extract_runs.create_and_poll(
      file={"url": "https://example.com/invoice.pdf"},
      config={
          "schema": schema,
          "advancedOptions": {"citationsEnabled": True},
      },
  )

  if run.status == "PROCESSED":
      data = run.output.value
  elif run.status == "FAILED":
      print(run.failure_reason, run.failure_message)
  ```

  ```typescript Extend (TypeScript) theme={"system"}
  import { ExtendClient } from "extend-ai";

  const client = new ExtendClient();

  const run = await client.extractRuns.createAndPoll({
    file: { url: "https://example.com/invoice.pdf" },
    config: {
      schema,
      advancedOptions: { citationsEnabled: true },
    },
  });

  if (run.status === "PROCESSED") {
    const data = run.output?.value;
  } else if (run.status === "FAILED") {
    console.error(run.failureReason, run.failureMessage);
  }
  ```
</CodeGroup>

<Warning>
  Citations are **opt-in** on Extend. Chunkr always returned them. If your application reads citations, set `advancedOptions.citationsEnabled: true` or they will be missing from the response. This also enables per-field `ocrConfidence`.
</Warning>

File inputs follow the same rules as Parse: `{"url": ...}` or `{"id": ...}` from `files.upload`. See [Migrating Parse](/pages/migrate-to-extend/parse#file-inputs).

***

## Schema changes

Chunkr accepted any JSON Schema generated by Pydantic or Zod. Extend validates schemas more strictly. Review your schemas against these rules before migrating:

| Rule | Example |
| :- | :- |
| Root must be `type: "object"` | Same as Chunkr |
| **Every primitive must be nullable** | `"type": ["string", "null"]`, not `"type": "string"` |
| Objects and arrays are **not** nullable | `"type": "object"`, never `["object", "null"]` |
| No `anyOf`, `oneOf`, `allOf`, `const`, or regex `pattern` | Replace unions with a single nullable type |
| Enums must include `null` and be strings only | `"enum": ["paid", "unpaid", null]` |
| Maximum nesting depth of 5 | Flatten deeply nested models |
| Property names: letters, numbers, `_`, `-` only | |

<Warning>
  **Pydantic gotcha.** `Optional[str]` generates `"anyOf": [{"type": "string"}, {"type": "null"}]`, which Extend rejects. Either write the schema as a plain dictionary (shown below) or post-process the generated schema to use `"type": ["string", "null"]`. The Extend TypeScript SDK accepts Zod schemas directly and handles this for you.
</Warning>

### Before and after

<CodeGroup>
  ```python Chunkr (Pydantic) theme={"system"}
  from typing import List, Optional
  from pydantic import BaseModel

  class LineItem(BaseModel):
      description: str
      quantity: float
      line_total: float

  class Invoice(BaseModel):
      invoice_number: str
      invoice_date: str
      vendor_name: Optional[str] = None
      line_items: List[LineItem]
      total_amount: float

  schema = Invoice.model_json_schema()
  ```

  ```python Extend (JSON Schema) theme={"system"}
  schema = {
      "type": "object",
      "properties": {
          "invoice_number": {
              "type": ["string", "null"],
              "description": "The invoice number",
          },
          "invoice_date": {
              "type": ["string", "null"],
              "extend:type": "date",  # Returned as yyyy-mm-dd
          },
          "vendor_name": {"type": ["string", "null"]},
          "line_items": {
              "type": "array",
              "items": {
                  "type": "object",
                  "properties": {
                      "description": {"type": ["string", "null"]},
                      "quantity": {"type": ["number", "null"]},
                      "line_total": {"type": ["number", "null"]},
                  },
              },
          },
          "total_amount": {"type": ["number", "null"]},
      },
  }
  ```

  ```typescript Extend (Zod) theme={"system"}
  import { extendDate, extendCurrency } from "extend-ai";
  import { z } from "zod";

  // Pass the Zod schema directly to config.schema — no toJSONSchema() needed
  const schema = z.object({
    invoice_number: z.string().nullable().describe("The invoice number"),
    invoice_date: extendDate(),
    vendor_name: z.string().nullable(),
    line_items: z.array(
      z.object({
        description: z.string().nullable(),
        quantity: z.number().nullable(),
        line_total: z.number().nullable(),
      })
    ),
    total_amount: extendCurrency(),
  });
  ```
</CodeGroup>

<Tip>
  Extend adds a few typed helpers worth adopting: `"extend:type": "date"` normalizes dates to `yyyy-mm-dd`, `"extend:type": "currency"` returns `{ amount, iso_4217_currency_code }`, and `"extend:type": "signature"` detects signature blocks. See the [schema reference](https://docs.extend.ai/extraction/schema).
</Tip>

Chunkr's `system_prompt` parameter maps to `config.extractionRules`, a plain-language string applied across the whole extraction:

```json theme={"system"}
{
  "config": {
    "extractionRules": "If multiple totals appear, use the grand total. Return all dates in ISO 8601 format."
  }
}
```

***

## Reading the output

Chunkr returned three parallel objects that mirrored your schema: `results`, `citations`, and `metrics`. Extend returns `value` (your data) and a flat `metadata` map keyed by field path.

<CodeGroup>
  ```json Chunkr theme={"system"}
  {
    "output": {
      "results": {
        "invoice_number": "INV-001",
        "line_items": [{ "description": "Widget A", "quantity": 2 }]
      },
      "citations": {
        "invoice_number": [
          {
            "citation_type": "Segment",
            "content": "Invoice # INV-001",
            "page_number": 1,
            "page_width": 792, "page_height": 612,
            "bboxes": [{ "left": 450, "top": 120, "width": 100, "height": 20 }]
          }
        ],
        "line_items": [ { "description": [ ... ], "quantity": [ ... ] } ]
      },
      "metrics": {
        "invoice_number": { "confidence": "High", "citation_status": "Created" },
        "line_items": [ { "description": { "confidence": "High", ... } } ]
      }
    }
  }
  ```

  ```json Extend theme={"system"}
  {
    "output": {
      "value": {
        "invoice_number": "INV-001",
        "line_items": [{ "description": "Widget A", "quantity": 2 }]
      },
      "metadata": {
        "invoice_number": {
          "ocrConfidence": 0.99,
          "citations": [
            {
              "fileId": "file_...",
              "page": { "number": 1, "width": 612, "height": 792 },
              "polygon": [
                { "x": 459.8, "y": 88.4 }, { "x": 545.2, "y": 88.4 },
                { "x": 545.2, "y": 101.6 }, { "x": 459.8, "y": 101.6 }
              ],
              "referenceText": "Invoice # INV-001"
            }
          ]
        },
        "line_items": { "ocrConfidence": 0.98 },
        "line_items[0]": { "ocrConfidence": 0.98 },
        "line_items[0].description": { "ocrConfidence": 0.95, "citations": [ ... ] },
        "line_items[0].quantity": { "ocrConfidence": 0.98, "citations": [ ... ] }
      }
    }
  }
  ```
</CodeGroup>

### Field paths

Chunkr's docs described field paths like `line_items[0].description` as a way to think about the mirrored structure. On Extend, those paths are literally the keys of `metadata`.

<CodeGroup>
  ```python Python theme={"system"}
  value = run.output.value
  metadata = run.output.metadata

  invoice_number = value["invoice_number"]
  invoice_meta = metadata["invoice_number"]

  for i, item in enumerate(value.get("line_items", [])):
      desc_meta = metadata.get(f"line_items[{i}].description")
      if desc_meta and desc_meta.ocr_confidence is not None and desc_meta.ocr_confidence < 0.8:
          print(f"Review line item {i}: {item['description']}")
  ```

  ```typescript TypeScript theme={"system"}
  const { value, metadata } = run.output!;

  const invoiceNumber = value.invoice_number;
  const invoiceMeta = metadata.invoice_number;

  value.line_items.forEach((item, i) => {
    const descMeta = metadata[`line_items[${i}].description`];
    if (descMeta?.ocrConfidence != null && descMeta.ocrConfidence < 0.8) {
      console.log(`Review line item ${i}: ${item.description}`);
    }
  });
  ```
</CodeGroup>

### Confidence

| Chunkr | Extend | Notes |
| :- | :- | :- |
| `metrics[path].confidence` = `"High"` / `"Low"` | `metadata[path].ocrConfidence` (0 to 1) | Requires `citationsEnabled`. `null` when word-level confidence is unavailable. |
| | `metadata[path].reviewAgentScore` (1 to 5) | Opt-in: `advancedOptions.reviewAgent.enabled`. A second model audits each field. |
| | `metadata[path].logprobsConfidence` | Being phased out. Don't build on it. |
| `metrics[path].citation_status` | Absent | A field with no `citations` array had no citation generated. |

A simple replacement for Chunkr's `Low` flag is a threshold on `ocrConfidence`, for example `< 0.8`. For a stronger signal, enable the Review Agent and route anything with `reviewAgentScore <= 3` to review. See [Extend's confidence guide](https://docs.extend.ai/extraction/confidence-scores).

### Citations

Chunkr citations carried `bboxes[]` in `{left, top, width, height}` pixels. Extend citations carry a `polygon[]` of `{x, y}` points. Reduce the polygon to a rectangle and normalize by the page size to get a Chunkr-style box:

<CodeGroup>
  ```python Python theme={"system"}
  def citation_to_box(citation):
      """Return a normalized {left, top, width, height} box from an Extend citation."""
      xs = [p.x for p in citation.polygon]
      ys = [p.y for p in citation.polygon]
      page = citation.page
      left, right, top, bottom = min(xs), max(xs), min(ys), max(ys)
      return {
          "page_number": page.number,
          "left": left / page.width,
          "top": top / page.height,
          "width": (right - left) / page.width,
          "height": (bottom - top) / page.height,
      }

  for citation in run.output.metadata["invoice_number"].citations or []:
      print(citation.reference_text, citation_to_box(citation))
  ```

  ```typescript TypeScript theme={"system"}
  import { Extend } from "extend-ai";

  function citationToBox(citation: Extend.Citation) {
    // Return a normalized {left, top, width, height} box from an Extend citation
    const xs = citation.polygon!.map((p) => p.x);
    const ys = citation.polygon!.map((p) => p.y);
    const { number, width, height } = citation.page;
    const [left, right] = [Math.min(...xs), Math.max(...xs)];
    const [top, bottom] = [Math.min(...ys), Math.max(...ys)];
    return {
      pageNumber: number,
      left: left / width,
      top: top / height,
      width: (right - left) / width,
      height: (bottom - top) / height,
    };
  }

  for (const citation of run.output!.metadata.invoice_number.citations ?? []) {
    console.log(citation.referenceText, citationToBox(citation));
  }
  ```
</CodeGroup>

| Chunkr citation field | Extend citation field |
| :- | :- |
| `content` | `referenceText` |
| `page_number` | `page.number` |
| `page_width` / `page_height` | `page.width` / `page.height` |
| `bboxes[]` | `polygon[]` |
| `citation_type` (`Segment` / `Word`) | `advancedOptions.citationMode` (`line`, `word`, `block`) set on the request |
| `segment_type`, `segment_id` | Not provided |
| `ss_ranges`, `ss_sheet_name` | Not provided in citations |

<Note>
  Chunkr returned segment-level citations always and word-level when available. On Extend you choose one granularity per run with `citationMode`. `"line"` is the default; `"word"` is closest to Chunkr's word citations; `"block"` is closest to segment citations.
</Note>

***

## Reusing a parsed document

On Chunkr you passed a parse `task_id` as the `file` to run several extractions over one parse. On Extend, pass the same `file.id` to each extract run. Extraction runs parse internally, and `parseRunId` on the response tells you which parse run was used.

<CodeGroup>
  ```python Python theme={"system"}
  with open("contract.pdf", "rb") as f:
      uploaded = client.files.upload(file=f)

  parties = client.extract_runs.create_and_poll(
      file={"id": uploaded.id}, config={"schema": parties_schema}
  )
  terms = client.extract_runs.create_and_poll(
      file={"id": uploaded.id}, config={"schema": terms_schema}
  )
  ```

  ```typescript TypeScript theme={"system"}
  const uploaded = await client.files.upload(createReadStream("contract.pdf"), {});

  const parties = await client.extractRuns.createAndPoll({
    file: { id: uploaded.id },
    config: { schema: partiesSchema },
  });
  const terms = await client.extractRuns.createAndPoll({
    file: { id: uploaded.id },
    config: { schema: termsSchema },
  });
  ```
</CodeGroup>

To tune the parse that runs under extraction (the equivalent of Chunkr's `parse_configuration` on an extract task), set `config.parseConfig` with any [Parse option](/pages/migrate-to-extend/parse#configuration-mapping).

***

<Tip>
  **Production pattern: saved extractors.** Instead of sending `config` on every call, create an Extractor once and reference it by ID. Extractors are versioned, so you can publish a schema change without redeploying code, and you can run [evaluation sets](https://docs.extend.ai/evaluation/overview) against them to measure accuracy on your own documents. See [Processors](https://docs.extend.ai/evaluation/processors).

  ```python theme={"system"}
  run = client.extract_runs.create_and_poll(
      file={"id": uploaded.id},
      extractor={"id": "ex_...", "version": "latest"},
  )
  ```
</Tip>

<Accordion title="Configuration mapping">
  | Chunkr | Extend |
  | :- | :- |
  | `schema` | `config.schema` |
  | `system_prompt` | `config.extractionRules` |
  | `parse_configuration` | `config.parseConfig` |
  | `expires_in` | `client.extract_runs.delete(id)` and `client.files.delete(id)` |
  | Always-on citations | `config.advancedOptions.citationsEnabled: true` |
  | | `config.baseProcessor`: `extraction_performance` (default), `extraction_light` (cheaper), `extraction_operator` (long, table-heavy docs) |
  | | `config.advancedOptions.pageRanges` to limit pages |
  | | `config.advancedOptions.reviewAgent.enabled` for a verification pass |

  Full reference: [Extend extraction configuration](https://docs.extend.ai/extraction/configuration).
</Accordion>

## Next Steps

<Columns cols={2}>
  <Card title="Polling, Webhooks & Production" href="/pages/migrate-to-extend/task-handling" icon="bell">
    Async runs, webhooks, retention, and limits.
  </Card>

  <Card title="FAQ" href="/pages/migrate-to-extend/faq" icon="circle-question">
    Accounts, deployment, legacy API, and feature differences.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.