docTR vs Docling on Receipts
Single-Pass Speed vs Document Pipeline (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party head-to-head benchmark · 2 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. Docling’s marketed strengths (layout analysis, table recognition, reading-order reconstruction, long documents) are out of scope here, not disproven — this benchmark was not designed to measure them. Cloud/API OCR services, other open-source engines (only these two are compared), fine-tuned models, and Docling’s native structured output (which the benchmark does not score) are out of scope. The complete 8-engine roundup lives on the OCR vs VLM comparison.
Range statement: every number on this page applies to receipts only — SROIE 2019 English receipts and CORD v2 Indonesian receipts. One hardware tier (RTX 4090 at $0.76/hr, price timestamped August 2026), one LLM postprocessor (deepseek-v4-flash at temperature 0), fixed model versions (docTR v1.0.1, Docling 2.119.0). Do not extrapolate these results to invoices, tables, or complex layouts — the benchmark measures receipt OCR and receipt-field extraction only, and Docling’s pipeline capabilities on structured documents are exactly what it does not measure. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.
The architecture tax on a simple receipt: Docling’s staged pipeline — layout boxes, table detection, reading-order reconstruction — buys little on a single-page English receipt, and the meter shows it. On the same 361 SROIE receipts, same RTX 4090, same protocol, Docling’s raw CER is 3.0× worse than docTR’s (0.5909 vs 0.1971), it runs 6.7× slower at p50 (732.0 ms vs 108.7 ms), and costs 8.3× more per 1,000 pages ($0.3978 vs $0.0479). The inversion that keeps this honest: despite far worse raw text, Docling’s regex field F1 on SROIE (0.2237) beats docTR’s (0.0766) by 2.9× — then an LLM postprocessor swings the ranking back to docTR (0.6171 vs 0.5685).
The trade, in one pair of numbers: docTR reads a receipt page at 108.7 ms p50 for $0.048 per 1,000 pages; Docling reads it at 732.0 ms p50 for $0.398 per 1,000 pages — same receipts, same test split, same GPU. Neither engine “wins”; this page measures whether the pipeline overhead earns its keep on a plain receipt. Here, it does not — and what Docling’s overhead does buy (layout structure, tables, reading order) is deliberately unmeasured by this benchmark, not disproven by it.
What Docling Is (and Isn’t): Single-Pass OCR vs a Parsing Pipeline
The two engines sit on opposite sides of a fundamental architectural divide, and that divide — not a code difference or a tuning difference — is the entire story of this page. docTR is a single-pass neural OCR engine: a detection stage localizes text bounding boxes and a recognition stage transcribes the characters inside them, composed into one OCR predictor whose front-to-back forward pass produces raw text lines. There is no layout model, no table parser, and no reading-order reconstruction — what is printed is what comes out, in the order the recognizer reads it. Docling is not an OCR engine and not a vision-language model; it is a document-parsing pipeline. Per its own technical report (cited for architecture context, not for any number on this page), it stages a sequence of models per page — layout analysis, table detection, reading-order inference — aggregates them, and assembles an intermediate document object before emitting text, which is why its output carries structure (labels, order, zones) that docTR’s lines do not.
Why this mechanism matters for a benchmark: every staged model in Docling’s chain exists to exploit layout structure — a table to parse, a two-column form, a reading path that is not lexical order. A plain English receipt has almost none of that: single column, a few zones, a mostly predictable top-to-bottom path, no tables. The staged machinery still runs on every page — that is why it is slower and costlier — but with no structure to exploit, the overhead cannot convert into better text. This page isolates exactly that cost and shows what it does — and does not — buy.
Character Accuracy: The Pipeline Tax on Raw Text
On SROIE 2019, the raw-text gap is not close: CER 0.1971 (docTR) vs 0.5909 (Docling) — a 3.0× penalty — and WER 0.3199 vs 0.7596. Character Error Rate measures insertions, deletions, and substitutions divided by ground-truth characters — a CER of 0.197 means ~19.7 misread characters per 100; Word Error Rate applies the same edit-distance logic at word granularity. Docling’s 0.5909 ranks 7th of 8 engines in the underlying run, ahead of only Unlimited-OCR (0.6552, summary_metrics.csv cer column, all sroie_2019 rows) — the docTR-vs-Surya2 sibling head-to-head documented docTR as one of the two best recognizers in this same benchmark, and this page shows the same engine at the other end of the text-accuracy table is a pipeline, not a weaker recognizer family.
Source: summary_metrics.csv — cer and wer columns, sroie_2019 rows. docTR cer 0.19707 / wer 0.31990; Docling cer 0.59092 / wer 0.75961. Lower is better. 361 samples per engine; both error_rate 0.0.
| Metric (SROIE 2019, n=361) | docTR | Docling | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.1971 | 0.5909 | summary_metrics.csv · cer, doctr/sroie_2019 and docling/sroie_2019 rows |
| Word Error Rate (WER) | 0.3199 | 0.7596 | summary_metrics.csv · wer, same rows |
| Error rate (failed pages) | 0.0 | 0.0 | summary_metrics.csv · error_rate, same rows |
Table: summary_metrics.csv — cer / wer / error_rate columns, sroie_2019 rows. Exact values: docTR cer 0.19707 / wer 0.31990; Docling cer 0.59092 / wer 0.75961. Docling’s SROIE CER is the second-worst of the eight engines in the underlying run (ahead of only Unlimited-OCR’s 0.6552) — raw character accuracy is where the pipeline tax shows up first.
The Inversion: Regex Field Extraction Flips the Outcome
Benchmark both engines’ text through the same fixed regex patterns on the four SROIE receipt fields (company, date, address, total) — the traditional OCR + rule-based key-information-extraction (KIE) approach — and the ranking flips: Docling extracts fields at 0.2237 field F1 to docTR’s 0.0766, a 2.9× advantage. These are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text — postprocessed, not native structured output by either engine, and Docling’s native document model is not scored here.
Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth — 1.0 means every receipt field perfectly recovered, 0 means nothing. The mechanism behind the flip is the same architecture difference that caused the CER gap, working in the opposite direction: Docling’s document model reorders text into a reading path and associates labels with values, so its emitted text is closer in shape to what the fixed patterns expect; docTR’s clean-but-raw line text — accurate by CER, but with original casing and separator noise and no label framing — defeats the patterns. docTR’s regex field F1 of 0.0766 is the worst of all eight engines in the underlying run despite its best-class CER (summary_metrics.csv field_f1_regex and cer columns, all sroie_2019 rows); Docling’s 0.2237 ranks sixth. The same decoupling the sibling head-to-head documented at the top of the accuracy ladder (docTR vs Surya2) recurs here at the bottom: text accuracy is not field accuracy.
Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| Regex postprocessing (SROIE 2019, n=361) | docTR | Docling | Source |
|---|---|---|---|
| Field-value F1 (regex) | 0.0766 | 0.2237 | field_method_comparison.csv · regex_field_value_f1, doctr/sroie_2019 and docling/sroie_2019 rows |
| Field-value accuracy (regex) | 0.0623 | 0.2043 | field_method_comparison.csv · regex_field_value_accuracy, same rows |
| Document-fields exact (regex) | 0.0000 | 0.0000 | field_method_comparison.csv · regex_document_fields_exact, same rows |
Table: field_method_comparison.csv — regex columns, sroie_2019 rows. These are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text, not native structured extraction. Neither engine gets all four fields exactly right through regex on any single SROIE receipt (0.0000, a literal zero recorded in the CSV). docTR’s regex field F1 of 0.0766 is the lowest of all eight engines in the underlying run.
LLM Postprocessing Restores the Ranking — Partially
Feed both engines’ text to an LLM postprocessor (deepseek-v4-flash at temperature 0) with a structured extraction prompt, and docTR regains the lead: field F1 0.6171 vs 0.5685 — a 0.049-point edge, small compared with the raw CER gap but not erased. The cleaner base text surfaces more recoverable field values; the LLM partially compensates for Docling’s layout artifacts but does not wipe them out.
This is the same convergence band seen across the full eight-engine benchmark — LLM postprocessing pulls healthy engines together because it understands semantics (numbers, dates, names) instead of matching character shapes — and the residual gap matters: docTR’s 0.6171 is the best LLM field F1 of all eight engines, while Docling’s 0.5685 ranks sixth (field_method_comparison.csv llm_field_value_f1, all sroie_2019 rows). The harsher bar — documents where all four fields match exactly — splits them 2.7×: docTR 0.1496 vs Docling 0.0554. Two costs come with the lever: an LLM call adds ~2.0–2.4 s of median latency per document on top of OCR time (1,996.3 ms for docTR’s text, 2,365.1 ms for Docling’s — API-incurred and identical in kind), and it cannot rescue text an engine fundamentally failed to read.
| LLM postprocessing (SROIE 2019, n=361) | docTR | Docling | Source |
|---|---|---|---|
| Field-value F1 (LLM) | 0.6171 | 0.5685 | field_method_comparison.csv · llm_field_value_f1, doctr/sroie_2019 and docling/sroie_2019 rows |
| Field-value accuracy (LLM) | 0.6170 | 0.5665 | field_method_comparison.csv · llm_field_value_accuracy, same rows |
| Document-fields exact (LLM) | 0.1496 | 0.0554 | field_method_comparison.csv · llm_document_fields_exact, same rows |
| Median LLM postprocess latency (ms) | 1,996.3 | 2,365.1 | field_method_comparison.csv · llm_median_latency_ms, same rows |
Table: field_method_comparison.csv — llm_* columns, sroie_2019 rows. LLM model: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms). Both rows completed with llm_ok_count 361.
The Operating Envelope: 6.7× Latency, 7.9× Throughput, 8.3× Cost
The pipeline tax is heaviest where throughput planning lives. On the same RTX 4090 at the same recorded $0.76/hr rate, docTR sustains 449.3 pages/min at 108.7 ms p50 per page for $0.048 per 1,000 pages; Docling sustains 56.7 pages/min at 732.0 ms p50 for $0.398 per 1,000 pages — a 6.7× latency gap, a 7.9× throughput gap, and an 8.3× cost gap. The tail is proportionally worse for the pipeline: p95 281.4 ms vs 3,239.8 ms, an 11.5× gap, because Docling’s staged models compound their worst-case timings page to page.
Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization — the price you would actually pay for the GPU time. Throughput is wall-clock pages per minute including that same initialization. Latency p50/p95 are steady-state per-page inference times measured warm-then-scored (model loading excluded). docTR is the fastest and the cheapest engine of all eight in the underlying run on SROIE; Docling, at 56.7 pages/min and $0.398 per 1,000 pages, sits in the bottom half of the operating-envelope table (summary_metrics.csv, latency_p50_ms / pages_per_minute / cost_per_1000_pages, all sroie_2019 rows).
Source: summary_metrics.csv — latency_p50_ms / latency_p95_ms columns, sroie_2019 rows. docTR p50 108.72 / p95 281.38; Docling p50 732.00 / p95 3239.79. Steady-state latency (warm_then_scored measurement mode, excludes model loading).
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. docTR 0.0479, Docling 0.3978. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026). docTR is the cheapest engine of the eight in the underlying run.
| Operating envelope (SROIE 2019, n=361) | docTR | Docling | Source |
|---|---|---|---|
| Latency p50 (ms) | 108.7 | 732.0 | summary_metrics.csv · latency_p50_ms, doctr/sroie_2019 and docling/sroie_2019 rows |
| Latency p95 (ms) | 281.4 | 3,239.8 | summary_metrics.csv · latency_p95_ms, same rows |
| Pages per minute (wall-clock) | 449.3 | 56.7 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $0.048 | $0.398 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. Both engines GPU (RTX 4090, $0.76/hr price timestamped in manifests); cost includes model init. Exact values: docTR p50 108.72 / p95 281.38 / 449.31 pg/min / $0.0479; Docling p50 732.00 / p95 3239.79 / 56.66 pg/min / $0.3978.
CORD (Indonesian Receipts): Both Collapse, docTR’s LLM Field Recovery Still Leads
Neither engine was trained predominantly on Indonesian receipts, so CORD v2 (100 samples, nested fields menu/sub_total/total) functions as a cross-language stress test — and both collapse on raw CER: 0.9101 (docTR) and 0.9219 (Docling), a language-mismatch wash. Per the benchmark protocol, CORD numbers are kept quarantined from the SROIE comparison — never merged into any ranking — because CORD’s ground-truth text embeds annotation structure, which inflates raw CER for every engine on top of the genuine language mismatch.
On the field metrics, Docling’s one preserved advantage narrows to near-nothing: through regex patterns, both engines recover almost no CORD fields (docTR 0.0000 — a literal zero in the CSV — vs Docling’s 0.0612, because the English-format patterns were never written for Indonesian text). The LLM postprocessor absorbs the language shock on both sides but keeps docTR ahead: field F1 0.5500 vs 0.4695. CORD is quoted here for language-robustness context; it is deliberately never pooled with the SROIE numbers into a single leaderboard.
| CORD v2, Indonesian receipts (n=100) | docTR | Docling | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.9101 | 0.9219 | summary_metrics.csv · cer, doctr/cord_v2 and docling/cord_v2 rows |
| Field-value F1 (regex) | 0.0000 | 0.0612 | field_method_comparison.csv · regex_field_value_f1, same rows |
| Field-value F1 (LLM) | 0.5500 | 0.4695 | field_method_comparison.csv · llm_field_value_f1, same rows |
| Cost per 1,000 pages | $0.094 | $0.538 | summary_metrics.csv · cost_per_1000_pages, same rows |
| Pages per minute (wall-clock) | 500.4 | 123.2 | summary_metrics.csv · pages_per_minute, same rows |
Table: summary_metrics.csv (cer / cost_per_1000_pages / pages_per_minute) and field_method_comparison.csv (field F1), cord_v2 rows. Do not merge CORD numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth, and the regex patterns were written for English formats. docTR’s CORD regex field F1 of 0.0000 is a literal zero recorded in the CSV, not a missing value.
Who Wins When: The Recap Grid
“Better” is workload-dependent, and this head-to-head splits the axes cleanly: on a plain English receipt, every speed/cost axis and raw-text axis favors docTR; the out-of-box regex field inversion favors Docling; an LLM postprocessor brings them back to a 0.049-point docTR edge; and the capabilities Docling exists for — layout, tables, reading order, long documents — are unmeasured here, not disproven.
Frequently Asked Questions
Is Docling more accurate than docTR on receipts?
No — on raw character accuracy, docling is 3.0× worse: SROIE CER 0.5909 vs 0.1971, and WER 0.7596 vs 0.3199 (summary_metrics.csv, cer / wer, sroie_2019 rows). Docling “wins” only one measured axis: out-of-box regex field extraction (0.2237 vs 0.0766 field F1) — and an LLM postprocessor swings that back to docTR (0.6171 vs 0.5685).
Why does Docling extract fields better with regex despite much worse raw text?
Because the two metrics grade different things, and Docling’s output shape happens to fit the patterns. Docling’s document pipeline reorders text into a reading path and associates labels with values, so its emitted text is structurally closer to what the fixed regex patterns expect; docTR emits clean raw line text that is accurate by CER but defeats the patterns (0.0766 field F1, the worst of the eight engines, against best-class CER). These are postprocessed_sroie_receipt_regex_* scores — OCR text run through fixed patterns — not native structured output (field_method_comparison.csv, regex_field_value_f1, sroie_2019 rows). The same text-accuracy ≠ field-accuracy decoupling appears across this benchmark.
Why is Docling so much slower and costlier per page?
Because it runs a staged document pipeline — layout analysis, table detection, reading-order reconstruction, and an intermediate document model — on every page, even when the page is a plain receipt with no structure to exploit. On SROIE that tax measures 6.7× at p50 (108.7 vs 732.0 ms), 11.5× at p95, 7.9× lower throughput (449.3 vs 56.7 pages/min), and 8.3× higher cost per 1,000 pages ($0.048 vs $0.398) — same GPU, same protocol (summary_metrics.csv, sroie_2019 rows).
Does an LLM postprocessor close the gap between docTR and Docling?
Mostly, but not fully: LLM postprocessed field F1 on SROIE lands at docTR 0.6171 vs Docling 0.5685 — a 0.049-point lead for docTR that survives the LLM’s partial compensation for Docling’s layout artifacts (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows). The cost of convergence is ~2.0–2.4 s of additional median LLM latency per document (llm_median_latency_ms, same rows).
Does this benchmark mean Docling is bad?
No — it means Docling’s strengths are not measured here. Docling is a document-parsing pipeline whose value proposition — layout structure, tables, reading order, forms, long documents — is exactly what a receipts-only benchmark cannot test. What this page shows is narrower: on a plain single-page receipt, the pipeline overhead does not earn its keep (3.0× worse CER, 8.3× cost), and its one measured advantage (2.9× regex field F1) is erased by an LLM postprocessor. The honest framing is scope, not verdict.
Why do both engines score so badly on CORD receipts?
Two compounding causes that the protocol keeps apart from the SROIE ranking: a genuine language mismatch (Indonesian receipts outside both engines’ training focus) and annotation-structure inflation inside CORD’s ground-truth text — CER lands at 0.9101 (docTR) and 0.9219 (Docling) (summary_metrics.csv, cer, cord_v2 rows). Under an LLM postprocessor docTR’s field F1 holds at 0.5500 vs Docling’s 0.4695 — the SROIE ordering, compressed. CORD rows are quoted and never pooled into a combined ranking.
Which engine should a receipt pipeline pick, docTR or Docling?
For high-volume receipt text with metered cost, docTR’s envelope is decisive: 108.7 ms p50, 449.3 pages/min, $0.048 per 1,000 pages — the fastest and cheapest engine in the underlying eight-engine run. If your pipeline consumes out-of-box structured text without any postprocessor, Docling’s regex-field edge (0.2237 vs 0.0766) is a real starting advantage. If LLM postprocessing is part of the design, docTR stays ahead by 0.049 and is cheaper to feed. If your workload is layout-heavy documents — tables, forms, long reports — this benchmark is not the right evidence for the decision; it measures receipts only (see Limitations).
Where do the numbers on this page come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, regex field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports a head-to-head slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Both engines saw the same images, same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). Both runs completed with error_rate 0.0 on both datasets (summary_metrics.csv error_rate column). The underlying run contains eight engines total; this page compares only the two named engines, with other engines quoted solely as ranking context. The complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.
Runtime Environment
- Hardware: both engines ran on the same NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026).
- Engines: out-of-the-box, no fine-tuning. Versions locked: docTR v1.0.1 (single-pass neural OCR — detection stage + recognition stage composed into one OCR predictor, GPU) and Docling 2.119.0 (document-parsing pipeline — layout analysis, table detection, reading-order reconstruction staged around an OCR core, GPU) — per the public repo model table (README.md) and run manifests.
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on both engines.
- Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
- Field postprocessing: SROIE regex field metrics are
postprocessed_sroie_receipt_regex_*(field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set. They measure OCR + downstream extraction, not native structured output by either model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended, and Docling’s native document model is not scored by this benchmark.
Metric Definitions
- CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better.
- WER (Word Error Rate): the same edit-distance calculation at word granularity.
- Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
- Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
- Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
- Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init. They measure different clocks.
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate, including model initialization.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the doctr and docling rows here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the doctr and docling rows here (and to all eight sroie_2019 rows in the ranking context).
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
- Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
- Auer et al., "Docling Technical Report" (2024). Architecture background only — describes Docling’s staged pipeline (layout analysis, table detection, reading-order inference, document assembly). No benchmark numbers on this page are taken from it.
Limitations
- Document scope — receipts only: SROIE + CORD. Nothing here measures the layout/table/reading-order/long-document handling that defines Docling’s value proposition; those capabilities are out of scope, not disproven. Do not use this page to conclude “Docling is bad.” It concludes: on a plain single-page receipt, the pipeline overhead does not earn its keep.
- Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; the gaps documented here (3.0× CER, 6.7× p50, 8.3× cost) are far beyond the noise band, but single-percent differences should be treated as noise, not engineering truth.
- Single GPU tier and single price: all numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
- Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the 0.049-point docTR lead may move at the margins. LLM latency (~1,996–2,365 ms median on SROIE, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of either engine’s own latency.
- Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes; Docling’s 2.9× regex edge is measured against this single fixed pattern set.
- Docling’s native output is unscored: Docling emits a structured document model, but the benchmark scores text + postprocessors, not native structured output. A benchmark variant scoring Docling’s native fields would be a different experiment; this page does not attempt it.
- CORD CER is not a per-model quality reading: CORD ground truth embeds annotation structure and neither engine was trained predominantly on Indonesian; CORD CER (~0.91–0.92) reflects language mismatch + ground-truth inflation. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
- Version pinning: results hold for docTR v1.0.1 and Docling 2.119.0 (August 2026). Newer releases of either engine may shift every number on this page.
Related references: docTR vs Surya2 Receipt Benchmark · PaddleOCR vs EasyOCR Receipt Benchmark · Traditional OCR vs Document Parsing VLMs · Rule-based extraction vs LLM extraction · Field-Level vs Character-Level Accuracy
Related reading: the accuracy gap between AI and traditional OCR · AI image extraction compared with traditional OCR · AI Document Extraction Pricing (2026)