PuffinParse
Docs Benchmark GitHub

Rust core · Python & TypeScript SDKs · gateway · open benchmark

One API for
every document parser.

Three modes — parse, ocr and extract — across 18 providers and 60 models. One request, one response shape, one set of errors, and an open benchmark that ranks every provider on accuracy, latency and cost.

pip install puffinparse
import puffinparse

doc = puffinparse.parse("invoice.pdf", model="reducto/standard")
print(doc.markdown, doc.pages[0].blocks[0].bbox)
print(doc.usage.pages, doc.cost_usd)
invoice.pdf reducto/standard
{
  "markdown": "# ACME Industries\n\n## Invoice …",
  "pages": [{"blocks": [
    {"type": "title",
     "bbox": [0.08, 0.07, 0.62, 0.11]},
    {"type": "table",
     "bbox": [0.08, 0.41, 0.92, 0.63]}
  ]}],
  "usage": {"pages": 2},
  "latency_ms": 2848,
  "cost_usd": 0.0300
}
markdowntyped blocks + boxesusage & cost

Any provider in, the same typed response out.

18providers
60models
3modes
1response shape

Switch in one line

Change the string. Nothing else moves.

Every provider has its own upload flow, polling loop, block vocabulary, coordinate system and billing unit. PuffinParse hides all of it. Pick a tab: only the highlighted model string changes.

import puffinparse

doc = puffinparse.parse("invoice.pdf", model="reducto/r-1")

doc.markdown          # same unified markdown
doc.pages[0].blocks   # same typed blocks, same normalised boxes
doc.usage.pages       # same billed page count
doc.cost_usd          # same field, priced from the model string
import puffinparse

doc = puffinparse.parse("invoice.pdf", model="extend/parse_performance")

doc.markdown          # same unified markdown
doc.pages[0].blocks   # same typed blocks, same normalised boxes
doc.usage.pages       # same billed page count
doc.cost_usd          # same field, priced from the model string
import puffinparse

doc = puffinparse.parse("invoice.pdf", model="mistral/ocr-latest")

doc.markdown          # same unified markdown
doc.pages[0].blocks   # same typed blocks, same normalised boxes
doc.usage.pages       # same billed page count
doc.cost_usd          # same field, priced from the model string
import puffinparse

doc = puffinparse.parse("invoice.pdf", model="gemini/2.5-flash")

doc.markdown          # same unified markdown
doc.pages[0].blocks   # same typed blocks, same normalised boxes
doc.usage.pages       # same billed page count
doc.cost_usd          # same field, priced from the model string

Providers are only interchangeable within a mode. A model that cannot serve the mode you asked for raises UnsupportedModelError before any network call — never a quietly different response shape.

Three products, not one

Document AI vendors sell three different things.

A call picks one mode, and the mode decides the response type. Providers are only interchangeable within a mode.

parse

Layout-aware markdown plus typed blocks with normalised boxes. For RAG chunks, tables and document structure.

{
  "markdown": "# Invoice …",
  "pages": [{"blocks": [
    {"type": "table",
     "bbox": [0.08, 0.41, 0.92, 0.63]}
  ]}]
}

ocr

Plain text in reading order with line and word boxes. For search indexes, redaction and overlays.

{
  "text": "ACME Industries …",
  "pages": [{
    "lines": [{"text": "Total due",
               "bbox": [0.6, 0.8, 0.8, 0.83]}],
    "words": [{"text": "Total", "bbox": […]}]
  }]
}

extract

A JSON object shaped by your schema, with per-field confidence and citations back to the page. For invoices, forms, anything with fields.

{
  "data": {"invoice_number": "INV-4182",
           "total": 1280.5},
  "fields": {"/total": {"confidence": 0.98}},
  "citations": {"/total": [{"page": 2,
                            "bbox": […]}]}
}

puffinparse.list_models("extract") or puffinparse providers --mode extract lists the models that serve a mode — it is a registry fact, not a guess.

Native-format compatibility

Already on Reducto, Extend or LlamaParse? Keep your code.

Ask for a vendor's own JSON shape and PuffinParse renders the unified response into it — whatever provider actually produced it. Key set, nesting, block vocabulary, coordinate convention and billed page count, all in the vendor's units.

Your parser is loyal to a vendor. Your code doesn't have to be.

How the compatibility layer works → including exactly what is guaranteed and what is rendered as null.

doc = puffinparse.parse("invoice.pdf",
                    model="gemini/2.5-flash",
                    output_format="reducto")
# -> Reducto's own parse JSON, from Gemini
doc.raw["result"]["chunks"][0]["blocks"][0]["bbox"]["left"]

Open benchmark

Pick a provider on evidence, not marketing.

Ground truth is exact by construction — the documents are rendered from the same source as the truth files. The metrics are deterministic text comparisons with no LLM judge, provider result caches are disabled, and every per-document output is committed next to the run so you can verify any number yourself.

#ModelOverallp50 latency$/1k pages
1llamaparse/cost_effective84.359,452 ms$3.75
2llamaparse/agentic83.4914,127 ms$12.50
3reducto/r-182.403,459 ms$10.00
4reducto/standard79.783,019 ms$15.00
5extend/parse_performance77.4321,965 ms$25.00
6extend/parse_light76.4632,038 ms$6.25

combined-v3 v3.0.0 · 199 documents (synthetic, ParseBench, olmOCR-bench and OmniDocBench pages, each scored by its own truth) · higher Overall is better

puffinparse bench run --dataset benchmark/datasets/combined-v2 \
    --models reducto/r-1 extend/parse_light llamaparse/cost_effective --dry-run
puffinparse bench report benchmark/results/2026-09-24-combined-v2.json

The results viewer opens every run document by document: the input, the ground truth, each model's raw output, a word-level diff and the commands to reproduce that exact score.

Results viewer → · Full leaderboard → · methodology and caveats → · raw result files on GitHub →

For agents

The docs are readable without a browser.

Every page on this site is also served as plain markdown, linked from the HTML with <link rel="alternate" type="text/markdown">. No JavaScript is needed to read anything. There is no MCP server yet — the markdown endpoints are the interface.