Surya2 vs Unlimited-OCR vs PaddleOCR-VL:
Receipt VLM Benchmark (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party three-way VLM benchmark · 3 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. The three engines’ marketed strengths (layout analysis, table recognition, formula parsing, and — for Unlimited-OCR — single-pass parsing of 40+ page documents) are not measured here. Cloud/API OCR services, the other five engines of the underlying run, and fine-tuned models are out of scope. The complete 8-engine roundup lives on text engines against document-understanding models.
Scope of every number on this page: receipts (SROIE 2019 English, CORD v2 Indonesian), one GPU tier (RTX 4090 at $0.76/hr), August 2026 model versions. Do not extrapolate these results to invoices, tables, or complex layouts — the benchmark measures receipt OCR and receipt-field extraction only. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.
Three document-parsing VLMs read the same 361 English receipts with raw character error rates that span 3.4× — SROIE 2019 CER 0.1915 (Surya2) vs 0.6552 (Unlimited-OCR), with PaddleOCR-VL between at 0.3370. That spread is mostly output convention, not reading ability: VLMs fold case, merge label/value lines, and reorder text, and CER counts every one of those normalizations as an error (the benchmark’s own decomposition attributes roughly a fifth of the SROIE CER budget to case substitutions alone). Put the three engines on the metrics their output is actually built for — field extraction — and the spread collapses: out-of-box regex field F1 0.3183–0.3376 across the trio, converging to 0.5921–0.6139 once an LLM postprocessor reads their text. Where the three genuinely separate is on Indonesian receipts (PaddleOCR-VL’s CORD regex field F1 of 0.3412 is the highest of all 8 engines in the benchmark) and on the operating envelope (3.8× latency and 5.2× cost gaps on identical hardware).
The trade, in one pair of numbers: PaddleOCR-VL reads a receipt page at 694.3 ms p50 for $0.205 per 1,000 pages; Surya2 reads it at 2,668.0 ms p50 for $1.061 per 1,000 pages — same receipts, same test split, same RTX 4090. The cheapest and slowest of the three are the same machine, and on receipts the character-level “quality” leader is the most expensive to run. None of the three “wins” everywhere; the point of this page is to show where each axis of the benchmark splits the field.
All three engines are document-parsing vision-language models (VLMs): neural models that read an entire document image and output understood text — case-folded, label/value pairs merged into single lines, rows reordered by reading order — rather than the raw character streams with original casing that traditional OCR engines (Tesseract, PaddleOCR, EasyOCR, docTR — the other engines in the underlying run) return. That output convention is what makes their field-extraction numbers strong out of the box and their raw character-error numbers misleading, as the next section shows. Within the VLM family, the three differ sharply in size and training target: Surya2 is a 650M-parameter, text-centric model tuned for clean full-page transcription (90+ languages); PaddleOCR-VL is a 0.9B compact generalist built for breadth across languages, tables, and formulas; Unlimited-OCR is oriented toward long-document and batch parsing (single-pass reading of 40+ page documents is its marketed core). One important caveat applies to every CER number below: Character Error Rate (CER) counts insertions, deletions, and substitutions against ground-truth characters, so it punishes exactly the normalizations VLMs are trained to perform. Field-level F1 is the fairer cross-VLM bar, and it is the spine of this page.
Why Not Use CER for VLMs: The 3.4× Spread Is Output Convention, Not Reading Ability
Read the raw CER column alone and Unlimited-OCR looks like a failed model (0.6552 on SROIE) while Surya2 looks world-class (0.1915, tied with the traditional docTR’s 0.1971 for the best raw CER in the 8-engine run). Both readings are artifacts of output style. The same Unlimited-OCR text that scores 0.6552 CER scores 0.4779 WER — its words survive while its characters look mangled, because case-folding replaces characters without breaking words. PaddleOCR-VL’s numbers invert the pattern: its CORD CER of 1.0805 is the worst of all 8 engines while its CORD field F1 of 0.3412 is the best of all 8 — the benchmark’s own CSV contradicts the CER ranking.
The mechanism has two layers. Layer 1 — normalization tax: document-parsing VLMs output “understood” text — TAN CHAY YEE becomes tan chay yee, INVOICE NO : PEGIV merges a label and value into one line. CER is exact-character matching, so every folded case and merged line is scored as an error even when the field value is correct. The benchmark’s error-decomposition analysis of the released SROIE predictions attributes roughly a fifth of the raw CER budget to case substitutions and roughly a tenth to line merges/drops; the three engines pay that tax at different rates — Unlimited-OCR’s heavy folding inflates its CER far beyond its WER, while PaddleOCR-VL’s line/label merging pushes its WER (0.6462) above its own CER (0.3370). Layer 2 — ground-truth structure inflation on CORD: CORD’s ground-truth text embeds annotation structure (menu entries, coordinates, field labels), so CER is systematically inflated for every engine on top of the genuine language mismatch — the all-engine CORD CER cluster of 0.90–1.08 (traditional Tesseract 0.9523, docTR 0.9101, and every VLM included) confirms the inflation is corpus-wide, not model-specific.
| Metric (SROIE 2019, n=361) | Surya2 | Unlimited-OCR | PaddleOCR-VL | Source |
|---|---|---|---|---|
| Character Error Rate (CER) | 0.1915 | 0.6552 | 0.3370 | summary_metrics.csv · cer, surya2/unlimited_ocr/paddleocr_vl_vllm sroie_2019 rows |
| Word Error Rate (WER) | 0.2735 | 0.4779 | 0.6462 | summary_metrics.csv · wer, same rows |
Table: summary_metrics.csv — cer and wer columns, sroie_2019 rows. Exact values: Surya2 cer 0.19147 / wer 0.27352; Unlimited-OCR cer 0.65524 / wer 0.47788; PaddleOCR-VL cer 0.33696 / wer 0.64623. Lower is better; all three runs completed with error_rate 0.0. Do not rank VLMs on CER: Unlimited-OCR’s CER/WER divergence (0.6552 vs 0.4779) and PaddleOCR-VL’s CORD CER-vs-field-F1 inversion (see below) are output-convention artifacts of the exact kind this benchmark’s protocol flags for VLM rows.
The consequence of Layers 1 and 2 is that every remaining section of this page compares the three VLMs on field-extraction F1 (the metrics their structured output feeds directly) and on the operating envelope (latency, throughput, cost) — and quotes CER only alongside its caveats. This is the benchmark’s protocol rule for document-parsing VLM rows, and it is the correct lens: a receipt pipeline consumes fields (company, date, address, total), not character streams.
Out-of-Box Field Extraction: Structured Output Is the VLM Family Trait
Run each engine’s raw text through the same fixed regex patterns on the four SROIE receipt fields (company, date, address, total) — the traditional OCR + rule-based key-information-extraction (KIE) approach — and the three VLMs land within a 0.02-point band: Unlimited-OCR 0.3376, PaddleOCR-VL 0.3368, Surya2 0.3183. All three rank in the top four of the eight-engine run; two of them beat the best traditional engine (PaddleOCR’s 0.3254), and the third trails it by 0.007 points. Their “understood text” reaches field consumers even without any LLM postprocessor — the family trait that raw-character OCR engines lack.
Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth: 1.0 means every receipt field perfectly recovered, 0 means nothing. The mechanism behind the VLM family’s edge is the output shape described above — the same case-folded, label-structured text that inflates CER happens to match extraction patterns. The “regex field extraction” columns are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: they measure OCR text + downstream rule-based extraction, not native structured output, and the same pattern set was applied to every engine. For contrast, the traditional engines’ regex field F1 lands at 0.0766 (docTR), 0.1477 (EasyOCR), 0.2237 (Docling), and 0.2335 (Tesseract) — six of the seven non-VLM engines sit below the trio’s lowest member.
Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| Regex postprocessing (SROIE 2019, n=361) | Surya2 | Unlimited-OCR | PaddleOCR-VL | Source |
|---|---|---|---|---|
| Field-value F1 (regex) | 0.3183 | 0.3376 | 0.3368 | field_method_comparison.csv · regex_field_value_f1, surya2/unlimited_ocr/paddleocr_vl_vllm sroie_2019 rows |
| Field-value accuracy (regex) | 0.2999 | 0.3089 | 0.3102 | field_method_comparison.csv · regex_field_value_accuracy, same rows |
| Document-fields exact (regex) | 0.0194 | 0.0028 | 0.0028 | field_method_comparison.csv · regex_document_fields_exact, same rows |
Table: field_method_comparison.csv — regex columns, sroie_2019 rows. These are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text (postprocessed, not native extraction). Surya2’s 0.3183 is the lowest of the trio but still ranks fourth of eight engines and sits 0.007 below the best traditional engine (PaddleOCR 0.3254, summary_metrics.csv field_f1_regex, paddleocr/sroie_2019 row).
The LLM Lever: The Three Engines Converge
Feed all three engines’ OCR text to an LLM postprocessor (deepseek-v4-flash, temperature 0) with a structured extraction prompt, and the out-of-box band tightens into a near-tie: Surya2 0.6139, Unlimited-OCR 0.6054, PaddleOCR-VL 0.5921 — a 0.022-point spread, fully inside the benchmark’s 0.57–0.62 convergence band for healthy engines. The postprocessor, not the VLM, becomes the deciding component.
This is the same convergence pattern the full 8-engine run exhibits: an LLM understands semantics (numbers, dates, names) instead of matching character shapes, so it absorbs most differences in upstream text quality — as long as the text is readable enough to work from. All three VLMs qualify; all three land within the band. The lever carries a cost: an LLM call adds roughly 1.9–2.3 s of median latency per document on top of OCR time (1,946.7 ms for PaddleOCR-VL’s text, 1,982.0 ms for Unlimited-OCR’s, 2,261.7 ms for Surya2’s — API-incurred and identical in kind), which favors asynchronous batch processing over synchronous per-page waits. Document-level exactness — the fraction of receipts where all four fields matched — stays low for all three (0.1219–0.1551), a reminder that per-field F1 is the meaningful operational number.
| LLM postprocessing (SROIE 2019, n=361) | Surya2 | Unlimited-OCR | PaddleOCR-VL | Source |
|---|---|---|---|---|
| Field-value F1 (LLM) | 0.6139 | 0.6054 | 0.5921 | field_method_comparison.csv · llm_field_value_f1, surya2/unlimited_ocr/paddleocr_vl_vllm sroie_2019 rows |
| Field-value accuracy (LLM) | 0.6136 | 0.6046 | 0.5852 | field_method_comparison.csv · llm_field_value_accuracy, same rows |
| Document-fields exact (LLM) | 0.1551 | 0.1302 | 0.1219 | field_method_comparison.csv · llm_document_fields_exact, same rows |
| Median LLM latency (ms) | 2,261.7 | 1,982.0 | 1,946.7 | field_method_comparison.csv · llm_median_latency_ms, same rows |
Table: field_method_comparison.csv — llm_* columns, sroie_2019 rows. LLM model: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms).
CORD (Indonesian Receipts): The Compact Generalist Wins
CORD v2 (100 Indonesian receipts, nested fields menu/sub_total/total) is the benchmark’s cross-language stress test — and it is where the three VLMs genuinely separate. Through the same English-format regex patterns, PaddleOCR-VL extracts Indonesian receipt fields at 0.3412 field F1 — the highest of all 8 engines in the entire benchmark — while Surya2 manages 0.2458 and Unlimited-OCR falls to 0.1079. The compact generalist’s training breadth shows exactly where the text-centric and long-document models lose ground.
All CORD CER values are quarantined by the benchmark protocol and never merged into any SROIE ranking: CORD’s ground truth embeds annotation structure (inflating raw CER for every engine on top of the genuine language mismatch — the all-engine CORD CER cluster of 0.90–1.08), and the regex patterns were written for English formats. The CORD comparison below is field-metrics only. Under an LLM postprocessor the language shock is absorbed as it was on SROIE: the trio re-converges to 0.4678–0.5203 field F1 (Surya2 0.5203, PaddleOCR-VL 0.5198, Unlimited-OCR 0.4678) — the postprocessor, not the engine, does the cross-language heavy lifting.
Source: summary_metrics.csv — field_f1_regex column, cord_v2 rows (0–1 stored decimals shown as %). PaddleOCR-VL 0.3412 is the maximum field_f1_regex across all 16 rows of the file; next-best on CORD regex is Surya2 0.2458.
| CORD v2, Indonesian receipts (n=100) | Surya2 | Unlimited-OCR | PaddleOCR-VL | Source |
|---|---|---|---|---|
| Field-value F1 (regex) | 0.2458 | 0.1079 | 0.3412 | summary_metrics.csv · field_f1_regex, surya2/unlimited_ocr/paddleocr_vl_vllm cord_v2 rows |
| Field-value F1 (LLM) | 0.5203 | 0.4678 | 0.5198 | field_method_comparison.csv · llm_field_value_f1, same rows |
| Character Error Rate (CER) — quarantined | 0.8959 | 0.9224 | 1.0805 | summary_metrics.csv · cer, same rows |
Table: summary_metrics.csv (field_f1_regex / cer) and field_method_comparison.csv (llm_field_value_f1), cord_v2 rows. Do not merge CORD numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth (all engines cluster at 0.90–1.08 — traditional Tesseract 0.9523, docTR 0.9101 included); PaddleOCR-VL’s CER of 1.0805 is the highest of the 8 engines precisely because its clean, normalized output is farthest from CORD’s structure-laden ground truth — while its regex field F1 is the benchmark’s best.
The Operating Envelope: 3.8× Latency, 5.2× Cost
Field accuracy converges; operating cost does not. On the same RTX 4090 at the same recorded $0.76/hr rate, PaddleOCR-VL sustains 68.2 pages/min at 694.3 ms p50 per page for $0.205 per 1,000 pages; Unlimited-OCR sits mid-envelope at 34.4 pages/min, 1,600.7 ms p50, $0.388 per 1,000 pages; Surya2 is the price-and-latency outlier at 12.1 pages/min, 2,668.0 ms p50, $1.061 per 1,000 pages. The fastest VLM is 3.8× faster and 5.2× cheaper than the slowest on identical hardware.
Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization — the price you would actually pay for the GPU time. Throughput is wall-clock pages per minute including that same initialization. Latency p50/p95 are steady-state per-page inference times measured warm-then-scored (model loading excluded); Surya2’s tail is proportionally worse — 5,872.2 ms p95 against PaddleOCR-VL’s 1,154.3 ms — because VLM prefill/decode spikes dominate the tail on first pages. Two numbers to read together rather than against each other: PaddleOCR-VL has the lowest p50 but is outrun on wall-clock throughput by Unlimited-OCR on CORD (73.99 vs 67.13 pages/min) — wall-clock figures include model initialization, and Unlimited-OCR’s vLLM batch handling is efficient enough to flip the order there.
Source: summary_metrics.csv — latency_p50_ms column, sroie_2019 rows. Surya2 2667.9800, Unlimited-OCR 1600.7459, PaddleOCR-VL 694.2519. Steady-state latency (warm_then_scored measurement mode).
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. Surya2 1.0609, Unlimited-OCR 0.3879, PaddleOCR-VL 0.2048. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026).
| Operating envelope (SROIE 2019, n=361) | Surya2 | Unlimited-OCR | PaddleOCR-VL | Source |
|---|---|---|---|---|
| Latency p50 (ms) | 2,668.0 | 1,600.7 | 694.3 | summary_metrics.csv · latency_p50_ms, surya2/unlimited_ocr/paddleocr_vl_vllm sroie_2019 rows |
| Latency p95 (ms) | 5,872.2 | 2,521.9 | 1,154.3 | summary_metrics.csv · latency_p95_ms, same rows |
| Pages per minute | 12.1 | 34.4 | 68.2 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $1.061 | $0.388 | $0.205 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. All three engines GPU (RTX 4090, $0.76/hr price timestamped in manifests); cost includes model init, not pure steady-state throughput. Exact values: Surya2 p50 2668.0 / p95 5872.2 / 12.1 pg/min / $1.0609; Unlimited-OCR p50 1600.7 / p95 2521.9 / 34.4 pg/min / $0.3879; PaddleOCR-VL p50 694.3 / p95 1154.3 / 68.2 pg/min / $0.2048.
Who Wins When: The Recap Grid
“Better” is workload-dependent, and among these three VLMs the axes split cleanly: field accuracy converges (regex and LLM), raw character text favors Surya2, cross-language fields favor PaddleOCR-VL, and every cost/latency/throughput axis favors PaddleOCR-VL with Unlimited-OCR in the middle. The honest takeaway is that on receipts, with any postprocessor in the pipeline, the VLM choice matters little — and without one, the cheap compact generalist beats the expensive specialists on the axes that usually matter.
Frequently Asked Questions
Which document-parsing VLM is most accurate on receipts?
On field extraction, it is a three-way near-tie on English receipts: SROIE regex field F1 0.3183–0.3376 and LLM field F1 0.5921–0.6139 across Surya2, Unlimited-OCR, and PaddleOCR-VL (field_method_comparison.csv, sroie_2019 rows). On Indonesian receipts the answer changes: PaddleOCR-VL’s CORD regex field F1 of 0.3412 is the best of all 8 engines in the benchmark (summary_metrics.csv, field_f1_regex, cord_v2 rows). Raw CER should not be used to rank VLMs — it counts output conventions (case folding, line merging) as errors (see the “Why Not Use CER” section above).
Why does Unlimited-OCR have the worst CER but the best regex field F1 on SROIE?
Because the two metrics grade different outputs. Unlimited-OCR’s heavy case-folding inflates character-level errors — its CER of 0.6552 vs WER of 0.4779 is the tell — while the same normalized text happens to match the fixed extraction patterns better than any other engine: SROIE regex field F1 0.3376, the top of the 8-engine run (summary_metrics.csv cer / wer, field_method_comparison.csv regex_field_value_f1, sroie_2019 rows).
Why is PaddleOCR-VL’s CORD CER the worst in the benchmark while its CORD field F1 is the best?
Because CORD’s ground truth embeds annotation structure and PaddleOCR-VL’s output is the cleanest and most normalized — the farthest from that structure-laden text, so its edit distance is the largest (CER 1.0805). The same output style feeds extraction patterns well: CORD regex field F1 0.3412, the best of all 8 engines (summary_metrics.csv, cer / field_f1_regex, cord_v2 rows). This inversion is the benchmark’s own demonstration that CORD CER is not a per-model quality reading.
How much faster and cheaper is PaddleOCR-VL than Surya2?
3.8× lower p50 latency (694.3 ms vs 2,668.0 ms), 5.6× higher throughput (68.2 vs 12.1 pages/min), and 5.2× lower cost per 1,000 pages ($0.205 vs $1.061) on the same RTX 4090 at $0.76/hr (summary_metrics.csv, latency_p50_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows).
Does adding an LLM postprocessor make the three VLMs equal?
Nearly — 0.5921 to 0.6139 LLM field F1 on SROIE, a 0.022-point spread inside the benchmark’s 0.57–0.62 convergence band (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows). The cost of that convergence is roughly 1.9–2.3 s of extra median LLM latency per document (llm_median_latency_ms, same rows), which favors asynchronous batch processing.
Which of the three VLMs should a receipt pipeline pick?
It depends on the axis your pipeline consumes. For native cross-language fields with no postprocessor, PaddleOCR-VL’s CORD regex F1 (0.3412) is the only useful level measured. For cheapest, fastest VLM serving, PaddleOCR-VL again (694.3 ms, $0.205/1K pages). For clean raw English text when cost and latency don’t gate, Surya2’s CER (0.1915) is the strongest. For a mid-volume balanced point, Unlimited-OCR (1,600.7 ms, $0.388/1K pages, best out-of-box SROIE field F1). With an LLM postprocessor in the pipeline, the choice matters little on receipts — all three land in the convergence band. These results hold for English and Indonesian receipts on one GPU tier in August 2026; any production decision should re-run on the target corpus (see Limitations).
Where do the numbers on this page come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports a three-way slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. All three engines saw the same images, the same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). All three runs completed with error_rate 0.0 (summary_metrics.csv error_rate column). The underlying run contains eight engines total; this page compares only the three named VLMs, and the complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.
Runtime Environment
- Hardware: all three engines ran on the same NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026).
- Engines: out-of-the-box, no fine-tuning. Versions locked: Surya2 (surya-ocr 0.22.1) and PaddleOCR-VL 1.6, both vLLM-served; Unlimited-OCR vLLM-served without a publicly pinned version (see Limitations) — per the public repo model table (README.md) and run manifests.
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on all three engines.
- Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
- Field postprocessing: SROIE regex field metrics are
postprocessed_sroie_receipt_regex_*(field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set. They measure OCR + downstream extraction, not native structured output by any model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended.
Metric Definitions
- CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better. Unfair to document-parsing VLMs: it scores case folding, label/value merging, and line reordering as errors even when field values are correct. Quoted on this page only alongside its caveats.
- WER (Word Error Rate): the same edit-distance calculation at word granularity. Where CER and WER diverge sharply (Unlimited-OCR: 0.6552 vs 0.4779), the gap marks where output normalization, not misreading, is doing the damage.
- Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
- Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
- Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
- Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init.
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the surya2, unlimited_ocr, and paddleocr_vl_vllm rows here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the three named-engine rows here.
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
- Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- CER unfairness to VLMs is the reason this page is built on field metrics: document-parsing VLMs fold case, merge label/value lines, and reorder text, so raw CER scores output conventions as errors — Unlimited-OCR’s CER (0.6552) vs WER (0.4779) divergence and PaddleOCR-VL’s CORD inversion (worst CER 1.0805, best field F1 0.3412) are both artifacts of that tax. Any comparison that ranks VLMs on CER — including on this page — should be treated as a measure of output style, not reading ability.
- Document scope — receipts only: SROIE + CORD. Nothing here measures the layout/table/formula/long-document handling where document-parsing VLMs claim their biggest advantages; all three engines’ marketed strengths (including Unlimited-OCR’s 40+ page single-pass parsing) are unmeasured. Do not use this page to conclude “VLM X wins on everything.”
- Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths (including the 0.022-point LLM-F1 spread) should be treated as noise, not engineering truth.
- Single GPU tier and single price: all numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
- Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the convergence ordering may move at the margins. LLM latency (~1.9–2.3 s median, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of any engine’s own latency.
- CORD quarantine: CORD ground truth embeds annotation structure and the regex patterns were written for English formats; CORD CER (0.90–1.08 across all engines) reflects language mismatch + ground-truth inflation, not per-model quality. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
- Version pinning: results hold for Surya2 0.22.1, PaddleOCR-VL 1.6, and Unlimited-OCR vLLM-served (August 2026). Unlimited-OCR has no publicly pinned version number in the benchmark’s published model table, so its row cannot be traced to an exact release; newer releases of any engine may shift every number on this page.
- Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes.
Related references: docTR vs Surya2 Receipt Benchmark · Traditional OCR vs Document Parsing VLMs · which extraction method wins on messy layouts · per-field scoring vs per-character scoring · Receipt OCR Accuracy
Related reading: AI OCR versus traditional OCR · image data extraction vs OCR engines · AI Document Extraction Pricing (2026)