Traditional OCR vs Document Parsing VLMsReceipt Benchmark Results (2026)

Last reviewed: 2026-08-18 · Run tier: official · First-party benchmark · 8 models × 2 receipt datasets

What this page covers: A first-party, reproducible benchmark comparing 4 traditional OCR engines (Tesseract, PaddleOCR, EasyOCR, docTR), 3 document-parsing vision-language models (Surya2, Unlimited-OCR, PaddleOCR-VL), and 1 pipeline parser (Docling) on two receipt datasets — SROIE 2019 English receipts (361 samples) and CORD v2 Indonesian receipts (100 samples). Metrics reported: character error rate (CER), word error rate (WER), field-extraction F1 under two postprocessing methods, p50/p95 latency, pages per minute, and cost per 1,000 pages. Every number traces to a published CSV row in the public OCR benchmark repository (ImageToTableai/benchmark-ocr) — reproducible experimental data, not an aggregation of third-party reports.
What this page does NOT cover: Any document type other than receipts — no invoices, forms, contracts, or long documents. Cloud/API OCR services, fine-tuned document-AI models, table/formula/layout accuracy, and full-text metrics outside CER/WER are out of scope. Results are further contextualized by the third-party receipt and document-type accuracy aggregations on Receipt OCR Accuracy and the shift in accuracy by document type.

Scope of every number on this page: receipts (SROIE 2019 English, CORD v2 Indonesian). Do not extrapolate these results to invoices, tables, or complex layouts — the benchmark measures receipt OCR and field extraction only. All figures come from the benchmark's results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.

Document-parsing VLMs are not naturally better than traditional OCR on receipts. On raw character accuracy (SROIE 2019), the best VLM (Surya2, CER 0.191) and the best traditional engine (docTR, CER 0.197) are statistically tied, with traditional PaddleOCR third at 0.204. The clear advantages of VLMs — layout, tables, formulas, long documents — simply do not show up on a single-page English receipt. What does separate the families on receipts is the operating envelope: traditional engines cost and run far less, and after an LLM postprocessing step, six of eight engines converge to a 0.57–0.62 field-F1 band.

The trade, in one pair of numbers: docTR processes a page at 109 ms p50 for $0.048 per 1,000 pages, while Surya2 takes 2,668 ms p50 at $1.061 per 1,000 pages on the same RTX 4090, same receipts, same test split — a 24.5× latency gap and a 22× cost gap. Which family "wins" depends entirely on which axis you care about; the point of this page is to show both axes from the same controlled run.

0.191 · 0.197
SROIE CER for the best document-parsing VLM (Surya2) vs the best traditional engine (docTR) — a statistical tie, not a VLM win (summary_metrics.csv, cer, surya2/sroie_2019 and doctr/sroie_2019 rows)
24.5×
Latency gap between docTR (108.7 ms) and Surya2 (2,668.0 ms) p50 per page on SROIE — 22× on cost and 37× on pages/min (summary_metrics.csv, latency_p50_ms / cost_per_1000_pages / pages_per_minute, same two rows)
0.57–0.62
LLM postprocessed (deepseek-v4-flash) field-F1 band on SROIE: 6 of 8 engines converge here (docling 0.569 … docTR 0.617); EasyOCR 0.372 and Tesseract 0.439 fall below (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows)

Character Error Rate (CER) measures what fraction of individual characters are misread — deletions, insertions, and substitutions divided by ground-truth characters. It is the classic OCR yardstick, and it is where the "VLM superiority" narrative collapses on English receipts.

On SROIE 2019, the two best text recognizers are one VLM and one traditional engine, separated by 0.006 points: Surya2 at 0.191 and docTR at 0.197, with PaddleOCR third (0.204). The three remaining VLMs — PaddleOCR-VL 0.337, Docling 0.591, Unlimited-OCR 0.655 — sit at or below traditional engines like EasyOCR (0.283) and Tesseract (0.335).

Character Accuracy by Model on SROIE (English Receipts)

SROIE 2019 CER by model: Surya2 0.191 and docTR 0.197 are tied at the top (lower is better). Traditional engines cluster 0.20–0.34; PaddleOCR-VL 0.337, Docling 0.591, Unlimited-OCR 0.655 trail.

Source: summary_metrics.csv — cer column, sroie_2019 rows (8 rows). Surya2 0.1915, docTR 0.1971, PaddleOCR 0.2045, EasyOCR 0.2833, Tesseract 0.3347, PaddleOCR-VL 0.3370, Docling 0.5909, Unlimited-OCR 0.6552. Lower is better. Tesseract is CPU-only.

ModelFamilyCERWERSource
Surya2Document-parsing VLM0.1910.274summary_metrics.csv · surya2/sroie_2019 row
docTRTraditional OCR0.1970.320summary_metrics.csv · doctr/sroie_2019 row
PaddleOCRTraditional OCR0.2040.326summary_metrics.csv · paddleocr/sroie_2019 row
EasyOCRTraditional OCR0.2830.616summary_metrics.csv · easyocr/sroie_2019 row
TesseractTraditional OCR (CPU)0.3350.559summary_metrics.csv · tesseract/sroie_2019 row
PaddleOCR-VLDocument-parsing VLM0.3370.646summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row
DoclingPipeline parser0.5910.760summary_metrics.csv · docling/sroie_2019 row
Unlimited-OCRDocument-parsing VLM0.6550.478summary_metrics.csv · unlimited_ocr/sroie_2019 row

Table: summary_metrics.csv — cer and wer columns, sroie_2019 rows, 361 samples each (error_rate 0.0 for all 8 models). CER = character error rate, WER = word error rate; lower is better. Exact values: Surya2 cer 0.19147 / wer 0.27352; docTR cer 0.19707 / wer 0.31990.

Word Error Rate (WER) tells the same story with a different granularity: it scores whole-word errors instead of characters. Surya2 leads WER at 0.274, docTR follows at 0.320. Note the outlier at the bottom: Unlimited-OCR has the worst CER (0.655) but a middle-of-pack WER (0.478) — its output is heavily case- and format-normalized (an output convention discussed in the methodology section), which inflates character-level edits even when words are largely intact.

Docling deserves a classification note before it appears in comparisons: it is neither a pure traditional OCR engine nor a VLM. Docling is a pipeline parser — a staged toolchain that runs layout analysis, table detection, and reading-order reconstruction around an OCR core. On a plain receipt that pipeline overhead buys little, which is part of why its raw CER (0.591 on SROIE) trails the single-pass engines.

Cost and Latency: The Traditional-Engine Advantage

If character accuracy decides nothing between the two families, cost and latency decide almost everything. On the identical test split, docTR sustains 449 pages/min at 108.7 ms p50 per page for $0.048 per 1,000 pages; Surya2 sustains 12 pages/min at 2,668 ms p50 for $1.061 per 1,000 pages — roughly 37× the throughput, 24.5× the per-page latency, and 22× the cost per thousand pages.

Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests) — the price you would actually pay for the GPU time, including model initialization. Tesseract is the special case: CPU-only, it has no GPU cost at all and still manages 78.6 pages/min on SROIE; its cost cell is empty in the CSV by design, not because it is free but because it consumes no billed GPU hours.

Median per-page latency (p50, ms) on SROIE 2019: docTR 109 ms. PaddleOCR 297, EasyOCR 414, Tesseract 671 (CPU), PaddleOCR-VL 694, Docling 732, Unlimited-OCR 1601, Surya2 2668. Vertical dashed bar at 1,000 ms marks the interactive-response threshold.

Source: summary_metrics.csv — latency_p50_ms column, sroie_2019 rows. docTR 108.7, PaddleOCR 297.0, EasyOCR 413.6, Tesseract 670.9 (CPU), PaddleOCR-VL 694.3, Docling 732.0, Unlimited-OCR 1600.7, Surya2 2668.0. Steady-state latency, warm-then-scored measurement mode (excludes model loading).

Cost per 1,000 pages on SROIE 2019 (RTX 4090 at $0.76/hr): docTR $0.048, EasyOCR $0.110, PaddleOCR-VL $0.205, PaddleOCR $0.221, Unlimited-OCR $0.388, Docling $0.398, Surya2 $1.061. Tesseract is CPU-only (no GPU cost, excluded).

Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. docTR 0.0479, EasyOCR 0.1098, PaddleOCR-VL 0.2048, PaddleOCR 0.2214, Unlimited-OCR 0.3879, Docling 0.3978, Surya2 1.0609. Tesseract CPU-only: cell empty in the CSV (no GPU cost); cost includes model init, not pure steady-state throughput.

ModelFamilyLatency p50 (ms)Latency p95 (ms)Pages/minCost / 1K pagesSource
docTRTraditional OCR108.7281.4449.3$0.048summary_metrics.csv · doctr/sroie_2019 row
PaddleOCRTraditional OCR297.03,331.479.7$0.221summary_metrics.csv · paddleocr/sroie_2019 row
EasyOCRTraditional OCR413.6960.4124.5$0.110summary_metrics.csv · easyocr/sroie_2019 row
TesseractTraditional OCR (CPU)670.91,507.078.6n/a (CPU)summary_metrics.csv · tesseract/sroie_2019 row
PaddleOCR-VLDocument-parsing VLM694.31,154.368.2$0.205summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row
DoclingPipeline parser732.03,239.856.7$0.398summary_metrics.csv · docling/sroie_2019 row
Unlimited-OCRDocument-parsing VLM1,600.72,521.934.4$0.388summary_metrics.csv · unlimited_ocr/sroie_2019 row
Surya2Document-parsing VLM2,668.05,872.212.1$1.061summary_metrics.csv · surya2/sroie_2019 row

Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. GPU runs on RTX 4090 ($0.76/hr, price timestamped in manifests); Tesseract ran CPU-only (empty cost cell, not zero). Throughput is wall-clock pages/min including model init.

The tail latency column matters if you care about worst-case behavior, not just medians. PaddleOCR's p95 of 3,331 ms and Docling's 3,240 ms are far from their p50 values — first-page effects and prefill spikes dominate the tail on GPU engines — while docTR's p95 (281 ms) stays tight. For interactive workloads (a user waiting on one page), that p95 spread is the difference between a 0.3-second wait and a 3+ second wait.

CORD (Indonesian Receipts): Language Mismatch and Ground-Truth Inflation

CORD v2 is an Indonesian-language receipt dataset with nested fields (menu, sub_total, total). None of the 8 engines was trained predominantly on Indonesian receipts, so CORD functions as a cross-language stress test — and every engine's CER collapses to 0.90–1.08. These numbers must be read with the caveat that CORD's ground-truth text embeds annotation structure, which inflates raw CER for every engine; CORD results are kept strictly separate from the SROIE ranking and are not mergeable into a single leaderboard.

Two distinct forces push CORD CER toward 1.0, and only one of them is the language itself. First, the language: English-trained engines genuinely misread Indonesian words — Indonesian names, street addresses, and currency formats (Rp) are outside their training distributions. Second, the ground truth: CORD's published text annotations embed the annotation structure (field labels with coordinates) rather than pure visible text, so raw CER measures edit distance against a structurally augmented string. The cleanest VLM examples are penalized hardest — PaddleOCR-VL at CER 1.080 is the extreme artifact of that mechanism, not a reading of its text quality.

The fair cross-family comparison on CORD is therefore the field metric, not CER (see the next section). What the CER columns still usefully show is that the language mismatch is real and universal across architectures — every family, traditional and VLM alike, lands in the same 0.90–1.08 band with no structural advantage for either.

ModelFamilyCORD CERSource
Surya2Document-parsing VLM0.896summary_metrics.csv · surya2/cord_v2 row
PaddleOCRTraditional OCR0.908summary_metrics.csv · paddleocr/cord_v2 row
docTRTraditional OCR0.910summary_metrics.csv · doctr/cord_v2 row
EasyOCRTraditional OCR0.918summary_metrics.csv · easyocr/cord_v2 row
DoclingPipeline parser0.922summary_metrics.csv · docling/cord_v2 row
Unlimited-OCRDocument-parsing VLM0.922summary_metrics.csv · unlimited_ocr/cord_v2 row
TesseractTraditional OCR (CPU)0.952summary_metrics.csv · tesseract/cord_v2 row
PaddleOCR-VLDocument-parsing VLM1.080summary_metrics.csv · paddleocr_vl_vllm/cord_v2 row

Table: summary_metrics.csv — cer column, cord_v2 rows, 100 samples each. Do not compare these numbers against SROIE in a combined ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth (methodology below). A CER above 1.0 (PaddleOCR-VL 1.0805) is an edit-distance artifact of that inflated ground truth.

Field F1: LLM Postprocessing Converges the Field

Character accuracy ranks engines; field extraction is what production users actually pay for. The benchmark extracts four receipt fields (company, date, address, total) from each engine's OCR text using two postprocessors — fixed regex patterns (the traditional OCR + rule-based KIE approach) and an LLM (deepseek-v4-flash) with a structured prompt. The result: the LLM almost erases the engine gap on SROIE, pulling six of eight engines into a 0.57–0.62 field-F1 band — where their regex results were spread across a 0.26-point range.

Field-value F1 is the harmonic mean of precision and recall over extracted field values, scored against ground truth — 1.0 means every field value perfectly extracted, 0 means nothing recovered. The regex columns use one fixed pattern set per dataset; the LLM columns use deepseek-v4-flash at temperature 0 for deterministic output (the llm_model column in the comparison CSV). The two metrics measure different pipelines and are never blended.

SROIE field F1 by postprocessing method: regex 7.7–33.8% vs LLM (deepseek-v4-flash) 37.2–61.7% across 8 engines. LLM converges 6 engines to 56.9–61.7%; EasyOCR 37.2% and Tesseract 43.9% fall below the band.

Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).

ModelFamilyregex field F1 (SROIE)LLM field F1 (SROIE)Source
docTRTraditional OCR0.0770.617field_method_comparison.csv · doctr/sroie_2019 row
Surya2Document-parsing VLM0.3180.614field_method_comparison.csv · surya2/sroie_2019 row
Unlimited-OCRDocument-parsing VLM0.3380.605field_method_comparison.csv · unlimited_ocr/sroie_2019 row
PaddleOCR-VLDocument-parsing VLM0.3370.592field_method_comparison.csv · paddleocr_vl_vllm/sroie_2019 row
PaddleOCRTraditional OCR0.3250.581field_method_comparison.csv · paddleocr/sroie_2019 row
DoclingPipeline parser0.2240.569field_method_comparison.csv · docling/sroie_2019 row
TesseractTraditional OCR (CPU)0.2330.439field_method_comparison.csv · tesseract/sroie_2019 row
EasyOCRTraditional OCR0.1480.372field_method_comparison.csv · easyocr/sroie_2019 row

Table: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1, sroie_2019 rows. LLM = deepseek-v4-flash postprocessing (llm_model column). docTR regex 0.0766 → LLM 0.6171 (8.1× lift); the 6 non-laggard engines span 0.5685–0.6171.

Two counter-intuitive results live in this table. First, docTR has the worst regex field F1 on SROIE (0.077) and the best LLM field F1 (0.617) — the same clean OCR text that regex mined into 7.7% of fields yielded 61.7% under the LLM. The postprocessor, not the OCR, was the bottleneck. Second, the two engines that fall outside the 0.57–0.62 band are exactly the two with degraded OCR bases: EasyOCR (0.372, SROIE CER 0.283) and Tesseract — whose CORD result (LLM field F1 0.163) shows that an LLM cannot extract fields from text it fundamentally cannot read (CORD CER 0.9523). The ceiling of any postprocessor is the OCR base quality beneath it.

On CORD, the LLM also absorbs part of the language shock: LLM field F1 holds at 0.47–0.55 for the healthy engines (PaddleOCR 0.553, docTR 0.550, PaddleOCR-VL 0.520, Surya2 0.520) even though regex collapses to near zero (docTR 0.0, EasyOCR 0.7%) — the patterns were written for English formats, and the penalty of "one more language" is paid almost entirely by the rules, not by the LLM (field_method_comparison.csv, llm_field_value_f1 / regex_field_value_f1, cord_v2 rows).

Why CER Understates Document-Parsing VLMs

CER compares VLM output and ground truth character-by-character, and VLMs are punished for two legitimate behaviors that are not recognition errors: case folding and label/value merging. The benchmark's CER decomposition on SROIE attributes roughly 18% of VLM CER to case-format differences (e.g., TAN CHAY YEE → tan chay yee) and roughly 10% to line-merging or dropped separator lines — with the field values themselves (company, total, date) actually correct (CER decomposition analysis recorded in the benchmark's protocol notes, SROIE runs).

The design tension is real and structural: traditional OCR engines output raw text with case intact, so they are optimized by design for CER scoring; document-parsing VLMs output "understood" text (normalized case, merged label-value pairs, reordered lines), which is closer to what a downstream system wants but farther from exact-character matches. This is why the page's headline uses CER only where it is a fair comparison among like outputs, and why the fair cross-family comparison lives in the field metrics — and why CORD CER (which additionally bears annotation-structure inflation) is quarantined to its own section. The broader point: a CER gap between families is not automatically an accuracy gap, and anyone comparing models across families should check what the CER is measuring before concluding one family "reads better."

How to Choose: Which Axis Matters for Your Workload

"Better" is meaningless without a workload. The benchmark's honest conclusion is that the two families win on different axes, and receipts specifically measure the axes where traditional engines win and the axes where postprocessing — not the engine family — decides field quality.

  1. Decide what your pipeline consumes: raw text or fields. If a human reads the text (search, display, audit), CER/WER is the honest metric — and traditional engines win or tie (SROIE CER: docTR 0.197 vs Surya2 0.191, summary_metrics.csv sroie_2019 rows). If a downstream system consumes fields, the postprocessor decides more than the engine: with regex, field F1 spans 0.077–0.338; with an LLM, six engines fit in 0.569–0.617 (field_method_comparison.csv sroie_2019 rows).
  2. If volume is high and cost is real, engineer around the traditional fast lane. docTR ran 449 pages/min at $0.048 per 1,000 pages (summary_metrics.csv doctr/sroie_2019: pages_per_minute 449.3, cost_per_1000_pages 0.0479). Tesseract adds zero GPU cost (CPU-only) at 78.6 pages/min. A VLM-based receipt line at Surya2's $1.061 per 1,000 pages costs roughly 22× more per page on the same hardware.
  3. If fields matter more than bytes, add LLM postprocessing rather than switching engines. The benchmark's largest single improvement was docTR's SROIE field F1 — from 0.077 (regex) to 0.617 (deepseek-v4-flash), an 8.1× lift from the same OCR text (field_method_comparison.csv doctr/sroie_2019 row). The LLM call adds ~1.8–2.4 s median per document (field_method_comparison.csv llm_median_latency_ms, all 16 rows) — suited to asynchronous batch processing, not synchronous per-page user waits.
  4. Budget for the ceiling your OCR sets. EasyOCR and Tesseract fall outside the LLM convergence band because their base text is weaker; Tesseract on CORD (LLM field F1 0.163 at CER 0.9523) is the hard proof that no postprocessor fixes unreadable text.
  5. Validate on your own documents before committing. These numbers came from one GPU tier (RTX 4090), two receipt datasets, and August 2026 model versions. Any architecture decision should re-run on your own corpus — the artifact set that produced this page exists precisely so that can happen.

Selection guidance is derived directly from the cited CSV rows; it is a data-driven reader aid, not a vendor endorsement. Your exact results vary with hardware, document mix, and model versions.

Frequently Asked Questions

Are document-parsing VLMs more accurate than traditional OCR on receipts?

Not on raw character accuracy — the best VLM and the best traditional engine are statistically tied on SROIE 2019 (Surya2 CER 0.191 vs docTR 0.197, summary_metrics.csv sroie_2019 rows), and PaddleOCR (0.204) is third. When field extraction is the goal, VLM or not, LLM postprocessing is the deciding factor (0.57–0.62 convergence band, field_method_comparison.csv).

When does traditional OCR make more sense than a document-parsing VLM?

When volume is high, cost is metered, or latency is interactive. On SROIE, docTR ran 449 pages/min at $0.048 per 1,000 pages and 108.7 ms p50; Surya2 ran 12 pages/min at $1.061 per 1,000 pages and 2,668 ms p50 (summary_metrics.csv, doctr and surya2 sroie_2019 rows). For a per-page interactive wait, the difference is 0.1 seconds vs 2.7 seconds.

Why do document-parsing VLMs sometimes score worse on CER than cheap OCR?

Because CER scores exact character matches, and VLMs are penalized for case folding and label/value merging that are output conventions, not misreads. The benchmark's CER decomposition on SROIE attributes roughly 18% of VLM CER to case-format differences and ~10% to line-merging/dropped separators — with the field values themselves often correct (see Methodology). CORD CER is additionally inflated by annotation structure inside its ground truth, which is why this page quarantines CORD CER from any ranking and uses field metrics for cross-family comparison.

How much does receipt OCR cost per page on an RTX 4090?

Between $0.048 (docTR) and $1.061 (Surya2) per 1,000 pages on an RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests (summary_metrics.csv cost_per_1000_pages, sroie_2019 rows). Tesseract is CPU-only and consumes no GPU hours. Cost includes model initialization, so per-page cost falls as batch size grows.

Why does every model score above 0.90 CER on CORD receipts?

Two compounding causes: genuine language mismatch (Indonesian receipts outside every engine's training focus) and annotation-structure inflation inside CORD's ground-truth text. Neither family escapes it — all 8 models land in the 0.90–1.08 band (summary_metrics.csv cer, cord_v2 rows). CORD is a language/layout-robustness stress set, kept separate from the SROIE ranking.

Will an LLM just fix my bad OCR output?

Only up to the quality of the base text. On SROIE the LLM lifted six engines into a 0.57–0.62 field-F1 band regardless of engine (field_method_comparison.csv), but Tesseract's CORD case shows the floor: at CER 0.9523 its LLM field F1 is 0.163 — an LLM cannot extract fields from text it cannot read.

Which is the fastest OCR for receipts?

docTR in this benchmark: 108.7 ms p50 per page and 449 pages/min on SROIE 2019 (summary_metrics.csv doctr/sroie_2019: latency_p50_ms, pages_per_minute). The slowest tested, Surya2, was 24.5× slower at p50 (2,668 ms) and 37× slower on throughput (12 pages/min).

Where do the numbers on this page come from?

Every figure is a row of the first-party benchmark's published CSVs — results/summary_metrics.csv (8 models × 2 datasets: CER/WER, field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.

Methodology & Sources

Protocol

This page reports the document-parsing comparison of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Every (model × dataset) pair reuses the same images, same ground truth, and same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). All 16 runs completed with error_rate 0.0 (summary_metrics.csv error_rate column).

Runtime Environment

  • Hardware: all GPU runs on an NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, with the price timestamp recorded in each run's redacted manifest (August 2026). Tesseract ran CPU-only and has no GPU cost (empty cost cell in the CSV).
  • Engines: all models run out-of-the-box, no fine-tuning. Versions locked per the run manifests.
  • LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows.
  • Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
  • Field postprocessing: SROIE field metrics are postprocessed_sroie_receipt_regex_* / LLM variants — i.e., fields extracted from OCR text by a fixed regex set or the LLM. They measure OCR + downstream extraction, not native structured output by the models.
ModelVersionType / Backend
Tesseract5.3.4Traditional OCR — CPU (no GPU cost)
PaddleOCR3.7.0Traditional OCR — GPU
EasyOCR1.7.2Traditional OCR — GPU
docTRv1.0.1Traditional OCR — GPU
Docling2.119.0Pipeline parser (layout + table + reading order) — GPU
Surya20.22.1Document-parsing VLM — vLLM-served
Unlimited-OCRvLLM-servedDocument-parsing VLM — vLLM-served
PaddleOCR-VL1.6Document-parsing VLM — vLLM-served

Versions as recorded on the benchmark's model table (README.md) and per-run redacted manifests (results/manifests/, one per published run, 16 total) — each manifest records run id, model version, runner-script hash, GPU/driver, torch/CUDA/Python versions, pip-freeze hash, cost metadata with price timestamp, and artifact hashes for reproducibility.

Metric Definitions

  • CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better. Sensitive to case and formatting conventions.
  • WER (Word Error Rate): the same edit-distance calculation at word granularity.
  • Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1.
  • Field-value F1 (LLM): the same metric on the LLM postprocessor's output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
  • Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init.
  • Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate; empty for CPU-only Tesseract.

Source List

  1. summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to a row here.
  2. field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to a row here.
  3. ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
  4. results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with the environment fingerprint, model version, cost metadata, and artifact hashes.
  5. Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
  6. Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).

Limitations

  • Document scope: Receipts only (SROIE + CORD). This benchmark measures nothing about layout/table/formula handling, long documents, or non-receipt fields — the document types where document-parsing VLMs claim their biggest advantages remain unmeasured here. Do not use this page to conclude "traditional OCR is better everywhere."
  • Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths should be treated as noise, not engineering truth.
  • Single GPU tier: all GPU numbers are from one RTX 4090 at $0.76/hr. Other GPUs, multi-GPU serving, or batch scheduling will shift latency, throughput, and cost.
  • CPU/GPU asymmetry: Tesseract (CPU) is compared against GPU-accelerated engines; its latency/time reflects CPU hardware while its cost advantage reflects zero GPU billing. This is marked on every relevant table but the asymmetry is inherent to the comparison.
  • LLM postprocessor is a single model: all LLM field rows use deepseek-v4-flash. A different LLM would produce different absolute F1; the convergence ordering may shift at the margins. LLM latency (~1.8–2.4 s median, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of the OCR engine's own latency.
  • Cost timestamp: the $0.76/hr GPU price was recorded in the run manifests in August 2026. GPU spot/on-demand prices change; re-derive costs at current rates before budgeting.
  • CORD CER is not a quality reading: CORD ground truth embeds annotation structure and the engines were not trained on Indonesian. CORD CER (0.90–1.08 across all 8 models) reflects language mismatch + ground-truth inflation, not per-model reading quality; CORD rows are intentionally not merged into any SROIE ranking.
  • Regex tuning: the regex pattern set was written once per dataset. A per-vendor, heavily tuned pattern library could score higher on its own formats — at the maintenance cost the LLM removes.
  • No cloud/API models: AWS Textract, Google Document AI, Azure AI Document Intelligence, and hosted VLM APIs (e.g., cloud OCR services) are not included; their latency and pricing models differ fundamentally from the local engines measured here.
  • Version pinning: results hold for the August 2026 model versions listed above; newer releases of any engine may shift results, and the p50 latencies of the two measurements with large p95 spikes (PaddleOCR, Docling) reflect prefill/first-page effects under this run's batch pattern.

Related references: the limits of regex for field extraction · field-level and character-level accuracy compared · Receipt OCR Accuracy · document-type OCR accuracy data

Related reading: AI OCR vs classic OCR accuracy · how AI vision extraction reads images differently from OCR · AI Document Extraction Pricing (2026)

📮 contact email: [email protected]