OCR Cost per 1,000 Pages
8 Open-Source Engines, Benchmarked (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party benchmark · 8 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no invoices, forms, contracts, or long documents. Cloud/API OCR services (AWS, Google, Azure and their per-call pricing), fine-tuned models, GPU procurement/amortization (buying hardware outright vs renting by the hour), storage and egress, and the LLM field-extraction dollar cost (token counts only) are out of scope. The full 8-engine accuracy/latency roundup lives on the eight-engine OCR vs VLM benchmark.
Range statement: all dollar figures on this page are for one GPU tier (RTX 4090) at one price timestamp (August 2026, $0.76/hr, recorded in the run manifests as price_recorded_at_utc 2026-08-13T08:00:00Z). Re-derive at current rates before budgeting. Receipt datasets only (SROIE 2019, CORD v2). Cost = wall-clock runtime × the hourly rate, including model initialization.
On the same GPU, the same receipts, and the same $0.76/hr billing basis, per-1,000-page OCR cost across 8 open-source engines spans 22.2× — from $0.0479 (docTR) to $1.0609 (Surya2) on SROIE 2019. Engine choice alone moves open-source OCR cost by more than an order of magnitude, and the ranking follows wall-clock time, not accuracy: the cheapest engine (docTR) has the second-best character error rate, while the priciest (Surya2) has the best.
The two numbers writers most often need: $0.048 per 1,000 pages for the cheapest engine measured (docTR, 449.3 pages/min) vs $1.061 per 1,000 pages for the priciest (Surya2, 12.1 pages/min) — same test split, same GPU tier, same price timestamp. Tesseract is CPU-only: it consumes no billed GPU hours and its cost cell is empty in the source CSV by design — not zero, and not free.
Cost per 1,000 pages is the billed GPU time for processing 1,000 pages at the recorded hourly rate — wall-clock run time × $0.76/hour, where wall-clock time includes model initialization. This page’s entire cost ranking is computed exactly that way; the worked arithmetic is shown in the How to Estimate Your Own Cost section and in each table’s source line.
Because the bill is metered by the hour, cost tracks time, not accuracy: an engine that reads a page in 109 ms costs ~25× less than one that reads it in 2,668 ms at the same rate. Every engine’s cost falls further as batch size grows, because the one-time model-init cost is amortized over more pages — the $ figures below reflect the benchmark’s run pattern (fixed test split, warm_then_scored measurement) and will not be the same as your own run’s cost.
SROIE 2019: Per-1,000-Page Cost Ranking
On 361 English receipts, the traditional two-stage OCR engines occupy the cheap end and the document-parsing VLMs the expensive end — but the spread inside each family is what surprises: PaddleOCR-VL, a compact 0.9B VLM, lands at $0.2048, within 4.3× of the cheapest engine, while its VLM sibling Surya2 costs 22.2× the cheapest — a 5.2× spread inside the VLM cluster alone.
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. docTR 0.0479, EasyOCR 0.1098, PaddleOCR-VL 0.2048, PaddleOCR 0.2214, Unlimited-OCR 0.3879, Docling 0.3978, Surya2 1.0609. Tesseract CPU-only: cell empty in the CSV (no billed GPU hours). Cost = wall-clock runtime × $0.76/hr (RunPod RTX 4090, price timestamped August 2026), including model init.
| Rank | Model | Type | Cost / 1K pages | Pages/min | Source |
|---|---|---|---|---|---|
| 1 | docTR | Traditional OCR (GPU) | $0.0479 | 449.3 | summary_metrics.csv · doctr/sroie_2019 row |
| 2 | EasyOCR | Traditional OCR (GPU) | $0.1098 | 124.5 | summary_metrics.csv · easyocr/sroie_2019 row |
| 3 | PaddleOCR-VL | Document-parsing VLM | $0.2048 | 68.2 | summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row |
| 4 | PaddleOCR | Traditional OCR (GPU) | $0.2214 | 79.7 | summary_metrics.csv · paddleocr/sroie_2019 row |
| 5 | Unlimited-OCR | Document-parsing VLM | $0.3879 | 34.4 | summary_metrics.csv · unlimited_ocr/sroie_2019 row |
| 6 | Docling | Pipeline parser | $0.3978 | 56.7 | summary_metrics.csv · docling/sroie_2019 row |
| 7 | Surya2 | Document-parsing VLM | $1.0609 | 12.1 | summary_metrics.csv · surya2/sroie_2019 row |
| — | Tesseract | Traditional OCR (CPU) | n/a (CPU-only, no GPU billing) | 78.6 | summary_metrics.csv · tesseract/sroie_2019 row |
Table: summary_metrics.csv — cost_per_1000_pages / pages_per_minute, sroie_2019 rows (361 samples each, error_rate 0.0). GPU runs at $0.76/hr (RTX 4090, price timestamped in manifests); Tesseract ran CPU-only (empty cost cell by design — no billed GPU hours, not a zero cost). Cost includes model init, so per-page cost falls with larger batches.
Cost Follows Throughput, Not Accuracy
Rank the engines by cost and by character accuracy and the two rankings barely agree. The best raw text quality on SROIE belongs to Surya2 (CER 0.1915) — the priciest engine at $1.0609 — while the second-best belongs to docTR (CER 0.1971) — the cheapest at $0.0479. Cost is a bill for time: Surya2’s 12.1 pages/min buys ~37× less throughput than docTR’s 449.3 pages/min at the same hourly rate.
The architecture-weight correlation is real but loose. As a class, the document-parsing VLMs (Surya2, Unlimited-OCR, PaddleOCR-VL) sit above the traditional two-stage OCR engines (docTR, EasyOCR, PaddleOCR), with the pipeline parser Docling between them. But inside each class the spread is wide — the VLM cluster spans 5.2× ($0.2048 to $1.0609) and the traditional cluster spans 4.6× ($0.0479 to $0.2214) — and PaddleOCR-VL, the smallest VLM in the run at 0.9B parameters, is within 4.3× of the cheapest engine while Surya2 is 22.2×. The practical takeaway: assume "more accurate = more expensive" is false until you measure it on your own documents; in this benchmark the relationship between cost rank and accuracy rank is effectively decoupled.
Source: summary_metrics.csv — pages_per_minute column, sroie_2019 rows. Wall-clock pages/min including model init. Tesseract ran CPU-only (78.6 pages/min on CPU hardware).
| Model | Type | CER (lower = better) | Latency p50 (ms) | Pages/min | Cost / 1K pages | Cost vs docTR | Source |
|---|---|---|---|---|---|---|---|
| docTR | Traditional OCR | 0.1971 | 108.7 | 449.3 | $0.0479 | 1.00× | summary_metrics.csv · doctr/sroie_2019 row |
| EasyOCR | Traditional OCR | 0.2833 | 413.6 | 124.5 | $0.1098 | 2.29× | summary_metrics.csv · easyocr/sroie_2019 row |
| PaddleOCR | Traditional OCR | 0.2045 | 297.0 | 79.7 | $0.2214 | 4.62× | summary_metrics.csv · paddleocr/sroie_2019 row |
| Tesseract | Traditional OCR (CPU) | 0.3347 | 670.9 | 78.6 | n/a (CPU) | n/a | summary_metrics.csv · tesseract/sroie_2019 row |
| PaddleOCR-VL | Document-parsing VLM | 0.3370 | 694.3 | 68.2 | $0.2048 | 4.28× | summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row |
| Docling | Pipeline parser | 0.5909 | 732.0 | 56.7 | $0.3978 | 8.31× | summary_metrics.csv · docling/sroie_2019 row |
| Unlimited-OCR | Document-parsing VLM | 0.6552 | 1,600.7 | 34.4 | $0.3879 | 8.10× | summary_metrics.csv · unlimited_ocr/sroie_2019 row |
| Surya2 | Document-parsing VLM | 0.1915 | 2,668.0 | 12.1 | $1.0609 | 22.16× | summary_metrics.csv · surya2/sroie_2019 row |
Table: summary_metrics.csv — cer / latency_p50_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. Cost ratios computed by simple division against docTR’s 0.0479 (e.g., 1.0609 / 0.0479 = 22.16). CER for Surya2 must be read with the case-folding caveat in the methodology (VLM outputs are case-normalized; CER overstates VLM error). Tesseract’s latency is CPU hardware; its cost cell is empty by design (no GPU billing).
The latency column adds the interactive-workload perspective: cost per volume and latency per page are two views of the same wall-clock fact. docTR’s p50 of 108.7 ms makes it both the cheapest per 1,000 pages and the only engine near interactive response; Surya2’s 2,668 ms p50 makes it both the priciest and a batch-only workload on receipts. The latency details are analyzed in the sibling 8-engine roundup.
CORD (Indonesian Receipts): Cost Is Dataset-Dependent
Swap the document set and the cost ranking shuffles — which proves per-1,000-page cost is not an engine constant. On CORD v2, EasyOCR becomes the cheapest engine ($0.0863) and docTR slips to second ($0.0939), while the anchors hold at both ends: docTR stays near the bottom and Surya2 stays the priciest ($1.1616). The CORD spread narrows to 13.5×.
The mechanism is throughput on the actual document set: CORD’s Indonesian receipts are shorter and less text-dense than SROIE’s English ones, so every engine’s pages-per-minute changes, and the cost per 1,000 pages follows. The two datasets are deliberately kept separate in this benchmark — CORD also carries an annotation-structure inflation in its ground truth that makes its CER readings unreliable (see methodology) — so treat the SROIE and CORD columns as two independent data points, not one ranking.
Source: summary_metrics.csv — cost_per_1000_pages column, cord_v2 rows. EasyOCR 0.0863, docTR 0.0939, Unlimited-OCR 0.2063, PaddleOCR-VL 0.2409, PaddleOCR 0.3419, Docling 0.5382, Surya2 1.1616. Tesseract CPU-only (empty cell). Same RTX 4090 @ $0.76/hr (price timestamped), cost incl. model init.
| Rank | Model | Type | Cost / 1K pages | Source |
|---|---|---|---|---|
| 1 | EasyOCR | Traditional OCR (GPU) | $0.0863 | summary_metrics.csv · easyocr/cord_v2 row |
| 2 | docTR | Traditional OCR (GPU) | $0.0939 | summary_metrics.csv · doctr/cord_v2 row |
| 3 | Unlimited-OCR | Document-parsing VLM | $0.2063 | summary_metrics.csv · unlimited_ocr/cord_v2 row |
| 4 | PaddleOCR-VL | Document-parsing VLM | $0.2409 | summary_metrics.csv · paddleocr_vl_vllm/cord_v2 row |
| 5 | PaddleOCR | Traditional OCR (GPU) | $0.3419 | summary_metrics.csv · paddleocr/cord_v2 row |
| 6 | Docling | Pipeline parser | $0.5382 | summary_metrics.csv · docling/cord_v2 row |
| 7 | Surya2 | Document-parsing VLM | $1.1616 | summary_metrics.csv · surya2/cord_v2 row |
| — | Tesseract | Traditional OCR (CPU) | n/a (CPU-only, no GPU billing) | summary_metrics.csv · tesseract/cord_v2 row |
Table: summary_metrics.csv — cost_per_1000_pages, cord_v2 rows (100 samples each). CORD is kept separate from the SROIE ranking (different language, different ground-truth structure); the point of the comparison is that per-1,000-page cost moves with the document set, not that one ranking "wins".
Monthly Volume Scenarios (Derived Estimates)
The per-1,000-page figures multiply directly into the volume math that budget writers need. At 100,000 pages/month on the SROIE cost basis, docTR’s GPU time is $4.79/month and Surya2’s is $106.09/month — a delta of roughly $101/month that widens to about $1,013/month at 1 million pages. These are arithmetic extensions of the measured per-1,000-page figures, not additional runs.
| Monthly volume | docTR GPU time (arithmetic) | Surya2 GPU time (arithmetic) | Delta | Basis |
|---|---|---|---|---|
| 10,000 pages | $0.48 (10 × $0.0479) | $10.61 (10 × $1.0609) | $10.13 | summary_metrics.csv cost_per_1000_pages, doctr / surya2 sroie_2019 rows; derived by simple multiplication — not a measured run |
| 100,000 pages | $4.79 (100 × $0.0479) | $106.09 (100 × $1.0609) | $101.30 | |
| 1,000,000 pages | $47.88 (1,000 × $0.0479) | $1,060.85 (1,000 × $1.0609) | $1,012.97 |
Table: Derived estimates on the SROIE 2019 cost basis — not measured runs. Each cell is a simple multiplication of the measured cost_per_1000_pages (docTR 0.0479, Surya2 1.0609; summary_metrics.csv sroie_2019 rows) by the volume in thousands, at the same $0.76/hr RTX 4090 rate, price timestamped August 2026. GPU time only — no CPU, storage, egress, orchestration, or LLM postprocessing. At 10 million pages/month Surya2 alone reaches roughly $10,609/month on this basis (10,000 × $1.0609).
Tesseract’s Real Cost Profile: No GPU Bill, But Not Free
Tesseract is the only CPU-only engine in the benchmark, and its cost cell is empty by design — it consumed no billed GPU hours, so there is nothing to bill at $0.76/hr. That is not the same as costing nothing: the CPU infrastructure it runs on (your own hardware or a rented CPU instance) is a real cost this benchmark does not quantify, and its field-recovery ceiling can push spend downstream.
On SROIE, Tesseract still sustained 78.6 pages/min on CPU — higher than four of the seven GPU engines (PaddleOCR-VL 68.2, Docling 56.7, Unlimited-OCR 34.4, Surya2 12.1). At low volume on already-provisioned CPU infrastructure, that makes it genuinely cost-effective: no GPU rental at all. The catch shows up when fields, not just text, are the deliverable: Tesseract’s weak base text caps what a downstream LLM can recover — its CORD LLM field F1 is 0.163 at a CORD CER of 0.9523 (field_method_comparison.csv, tesseract/cord_v2 row) — so the GPU savings can be offset by postprocessing and correction spend elsewhere. The CPU-vs-GPU tradeoff is the subject of the dedicated Tesseract vs PaddleOCR head-to-head; the postprocessor ceiling is covered in the rule-based vs LLM extraction comparison.
Pipeline Cost ≠ Engine Cost: The LLM Postprocessor Boundary
Every number on this page is the OCR engine’s GPU bill and nothing else. Production extraction pipelines commonly add an LLM field-extraction pass on top of the OCR text — the benchmark ran one (deepseek-v4-flash) across all 16 runs — and that pass is a separate, per-token API cost that does not appear in any engine figure here.
The comparison CSV records the token counts for the postprocessing pass on SROIE — for example, docTR’s OCR text cost 151,131 prompt + 25,377 completion tokens to extract the four receipt fields — and those token counts are the basis for estimating the additional spend. This page deliberately does not convert tokens to dollars: LLM pricing varies by provider, plan, and model, and any dollar figure would immediately go stale. Engine cost and pipeline cost are two lines in the budget; the LLM line is a function of your prompt design and provider, not of the OCR engine.
| Engine (OCR text source) | LLM prompt tokens (SROIE) | LLM completion tokens (SROIE) | Source |
|---|---|---|---|
| Tesseract | 128,285 | 24,561 | field_method_comparison.csv · tesseract/sroie_2019 row |
| PaddleOCR | 134,746 | 26,059 | field_method_comparison.csv · paddleocr/sroie_2019 row |
| EasyOCR | 138,689 | 25,607 | field_method_comparison.csv · easyocr/sroie_2019 row |
| docTR | 151,131 | 25,377 | field_method_comparison.csv · doctr/sroie_2019 row |
| Surya2 | 130,262 | 26,372 | field_method_comparison.csv · surya2/sroie_2019 row |
| Docling | 157,191 | 26,087 | field_method_comparison.csv · docling/sroie_2019 row |
| Unlimited-OCR | 185,705 | 25,876 | field_method_comparison.csv · unlimited_ocr/sroie_2019 row |
| PaddleOCR-VL | 148,973 | 26,301 | field_method_comparison.csv · paddleocr_vl_vllm/sroie_2019 row |
Table: field_method_comparison.csv — llm_prompt_tokens / llm_completion_tokens, sroie_2019 rows; llm_model = deepseek-v4-flash. Token counts cover one field-extraction pass (four receipt fields) over the 361-page SROIE test split. These are the basis for estimating LLM postprocessing spend; no dollar conversion is given because LLM pricing varies by provider and plan.
How to Estimate Your Own Cost
You do not need to re-run an 8-engine benchmark to get a defensible cost estimate for your own workload. The benchmark’s method — wall-clock time × your hourly rate — reproduces with four steps and one worked example.
- Pick your engine and its measured throughput. Use the pages-per-minute column for a same-document-type proxy (e.g., docTR 449.3 or Surya2 12.1 pages/min on SROIE; summary_metrics.csv sroie_2019 rows). For your own document mix, measure your own throughput on a small sample — the ranking above shows per-1,000-page cost is dataset-dependent.
- Convert volume to wall-clock hours.
hours = (model init + N / pages_per_minute) / 60for N pages. Model initialization is paid once per process/batch, so it must appear in the numerator. - Multiply by your hourly rate.
cost = hours × rate. The benchmark’s rate was $0.76/hr (RunPod RTX 4090, price timestamped August 2026). Worked example from the benchmark itself: the docTR SROIE run took 81,867 ms wall-clock for 361 pages (performance.run_wall_time_msin its redacted manifest) → 0.0227 hr × $0.76 = $0.0173 for the run → × 1,000 / 361 = $0.0479 per 1,000 pages, matching the CSV row exactly. - Account for initialization amortization and batch size. The 81,867 ms above includes init for a 361-page batch; at 1,000,000 pages the same init is diluted ~2,770× and per-page cost approaches pure steady-state throughput. Small batches pay the init cost repeatedly — a 10-page batch pays the same init as a 10,000-page run, so per-1,000-page cost rises sharply at small batch sizes. If your workload is bursty, either size batches up or accept init-heavy economics.
- Add pipeline costs on separate lines. LLM field extraction bills per token (token counts in the comparison CSV), storage and egress bill per byte, and Tesseract-style CPU-only stacks bill for CPU time — none of those are inside the per-1,000-page figures on this page.
This is a back-of-the-envelope method derived from the benchmark’s own cost basis (wall-clock × rate, incl. model init); it is not financial advice and your exact numbers vary with hardware, document mix, batch sizes, and utilization. The $0.76/hr rate is a timestamped August 2026 on-demand price — re-derive at current rates.
Frequently Asked Questions
How much does OCR cost per 1,000 pages?
Between $0.048 (docTR) and $1.061 (Surya2) per 1,000 pages on SROIE 2019, measured on an RTX 4090 at $0.76/hr with the price timestamped August 2026 (summary_metrics.csv cost_per_1000_pages, sroie_2019 rows). Cost includes model initialization, so per-page cost falls as batch size grows. On CORD v2 the same engines span $0.086 (EasyOCR) to $1.162 (Surya2).
Which open-source OCR engine is the cheapest to run?
docTR was the cheapest GPU engine measured at $0.0479 per 1,000 pages on SROIE (449.3 pages/min, summary_metrics.csv doctr/sroie_2019 row) — and the dataset-dependence caveat matters: on CORD v2, EasyOCR ($0.0863) edged out docTR ($0.0939). The answer to "cheapest" depends on your document set; the anchors (docTR near-low, Surya2 high) held across both.
Why is the fastest engine also the cheapest?
Because the bill is wall-clock hours at $0.76/hr: the engine that finishes a page in 109 ms (docTR) pays ~1/25 of the hours that a 2,668 ms engine (Surya2) pays per page (summary_metrics.csv latency_p50_ms / cost_per_1000_pages, sroie_2019 rows). Under metered GPU billing, throughput is cost — which is why the cost ranking and the throughput ranking are near-mirror images.
Is Tesseract OCR free?
No — Tesseract is CPU-only, so it consumes no billed GPU hours and its cost cell is empty in the benchmark CSV by design, but that is not a zero: the CPU infrastructure it runs on is a real cost, and its field-recovery ceiling (CORD LLM field F1 0.163, field_method_comparison.csv tesseract/cord_v2 row) can push spend into downstream postprocessing. At low volume on already-provisioned CPU hardware it can be genuinely cost-effective — 78.6 pages/min on SROIE, faster than four of the seven GPU engines — but "no GPU bill" and "free" are different statements.
Why is Surya2 so expensive per 1,000 pages?
Because it is the slowest engine in the benchmark: 12.1 pages/min on SROIE means the most wall-clock hours per 1,000 pages at the same $0.76/hr rate (summary_metrics.csv pages_per_minute / cost_per_1000_pages, surya2/sroie_2019 row). Notably it also has the best character accuracy (CER 0.1915) — the benchmark’s clearest example that cost follows time, not accuracy.
Does OCR cost per page drop as volume grows?
Yes, up to a floor. Cost here includes model initialization, which is paid once per process/batch; the benchmark’s docTR run paid init inside 81,867 ms for 361 pages (manifest performance.run_wall_time_ms), so at 1 million pages that init is diluted to near zero and cost approaches pure steady-state throughput. The floor is the steady-state cost itself — docTR’s $0.0479/1K on SROIE is already near its floor; Surya2’s $1.0609 reflects genuinely slow steady-state inference, not just init overhead.
What is not included in these per-1,000-page figures?
LLM postprocessing (a separate per-token API cost; token counts in field_method_comparison.csv), CPU/infrastructure for Tesseract-style stacks, storage and egress, orchestration, GPU utilization gaps, and cloud/API OCR services — none of which are in the engine GPU bill these figures represent. Cloud APIs also bill per call with feature-dependent pricing, a different cost model than rented-GPU wall-clock time; they are not benchmarked on this page.
Where do these numbers come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (cost per 1,000 pages, throughput, latency, accuracy) and results/field_method_comparison.csv (LLM postprocessor tokens and field F1) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run recording the $0.76/hr rate, the August 2026 price timestamp, model versions, and environment hashes.
Methodology & Sources
Protocol
This page reports the cost dimension of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Every (engine × dataset) pair reused the same images, the same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass). All 16 runs completed with error_rate 0.0 (summary_metrics.csv error_rate column).
Runtime Environment and Cost Basis
- Hardware: all GPU runs on an NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, with the price timestamp (
price_recorded_at_utc 2026-08-13T08:00:00Z) recorded in each run’s redacted manifest. Tesseract ran CPU-only and has no GPU cost (empty cost cell in the CSV — by design, not a zero). - Cost formula: cost per 1,000 pages = wall-clock runtime × $0.76/hr × (1,000 / pages processed), including model initialization. Verified by the worked docTR example in the How to Estimate section (81,867 ms → $0.0479/1K).
- Engines: all models run out-of-the-box, no fine-tuning. Versions locked per the run manifests (Tesseract 5.3.4, PaddleOCR 3.7.0, EasyOCR 1.7.2, docTR v1.0.1, Docling 2.119.0, Surya2 0.22.1, Unlimited-OCR vLLM-served, PaddleOCR-VL 1.6).
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 (the llm_model column in field_method_comparison.csv); its token counts are reported as the basis for separate pipeline cost — the OCR engine cost figures never include it.
- Field postprocessing: SROIE field metrics are
postprocessed_sroie_receipt_regex_*/ LLM variants — fields extracted from OCR text, not native structured output. - CORD caveat: CORD ground-truth text embeds annotation structure, which inflates raw CER for every engine; CORD rows are therefore kept separate from SROIE rankings. Cost figures (wall-clock-based) are not affected by the CER caveat, but both datasets are still receipts only.
Metric Definitions
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate, wall-clock including model initialization. Empty for CPU-only Tesseract.
- Pages per minute: wall-clock throughput including model init.
- Latency p50/p95: steady-state per-page inference time (warm-then-scored, excludes model loading).
- CER/WER: edit distance (insertions + deletions + substitutions) over ground-truth characters/words. Sensitive to case and formatting conventions — VLM outputs are case-normalized, so CER overstates VLM error (see the roundup page).
- Field-value F1: precision/recall harmonic mean over extracted field values, per postprocessor (regex or LLM).
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 engines × 2 receipt datasets (sroie_2019, cord_v2). Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every cost, throughput, latency, and CER figure on this page traces to a row here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/LLM field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, llm_prompt_tokens, llm_completion_tokens. Every token count and LLM field-F1 figure traces to a row here.
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with the environment fingerprint, model versions, cost metadata (
gpu_hourly_usd,price_recorded_at_utc), wall-clock time, and artifact hashes. - Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- Single GPU tier and single price timestamp: all GPU figures are from one RTX 4090 at $0.76/hr, price recorded August 2026. GPU spot/on-demand prices change — re-derive at current rates; other GPUs, multi-GPU serving, and batch scheduling shift throughput and cost.
- Receipts only: SROIE (English) and CORD (Indonesian) receipts. Cost, throughput, and accuracy on invoices, forms, contracts, or long documents are unmeasured; the rankings above are not generalizable beyond receipts.
- Cost includes model init — batch-size dependent: figures reflect the benchmark’s run pattern (361/100-page fixed splits, warm-then-scored). Smaller batches pay init repeatedly and cost more per 1,000 pages; larger batches approach the steady-state floor.
- CPU/GPU asymmetry: Tesseract (CPU) is compared against GPU-accelerated engines. Its cost advantage reflects zero GPU billing, not zero cost — CPU infrastructure, power, and staff time are not quantified, and its field-recovery ceiling (CORD LLM field F1 0.163) can shift spend into postprocessing.
- No cloud/API models: AWS Textract, Google Document AI, Azure AI Document Intelligence, and hosted OCR/VLM APIs are not included; their per-call, feature-metered pricing differs fundamentally from rented-GPU wall-clock billing and no comparison is implied.
- Utilization and idle time not modeled: figures assume the GPU is billed for the run’s wall-clock and is otherwise unused; real deployments with idle, multi-tenant, or underutilized GPUs have different effective costs.
- LLM postprocessing cost is not converted to dollars: token counts (field_method_comparison.csv) are the basis; LLM pricing varies by provider and plan and is intentionally left to the reader.
- Derived scenarios are not measured: the monthly volume table is simple arithmetic on the SROIE cost basis, labeled as derived estimates — not additional benchmark runs.
- Sample size and version pinning: 361 + 100 samples; results hold for the August 2026 model versions listed above. Newer engine releases may shift cost/throughput; single-digit-percent differences should be treated as noise.
Related references: Traditional OCR vs Document Parsing VLMs · docTR vs Surya2: CER Tie, Cost Gap · Tesseract vs PaddleOCR: CPU Cost Profile · how regex and LLMs extract fields · Document Processing Cost Breakdown
Related reading: AI Document Extraction Pricing (2026) · field-level accuracy in AI OCR vs traditional OCR · how AI vision extraction reads images differently from OCR