Receipt OCR Benchmark Results:Accuracy, Latency, and Cost Across 5 Open-Source OCR Models (2026)

Last reviewed: 2026-08-12 · Run tier: official · 10 runs (5 models × 2 datasets) · 461 receipts

What this page covers: First-party benchmark results from independent runs conducted by ImageToTable.ai using publicly available datasets and open-source models — 5 models on 461 receipt images across 2 public test sets. All prediction artifacts, metrics, and manifests are available for audit. This is NOT a third-party data aggregation — every number comes from reproducible benchmark runs on RTX 4090 hardware under a single measurement protocol.
What this page does NOT cover: Third-party aggregations of receipt OCR accuracy (see Receipt OCR Accuracy), cloud/API models (AWS Textract, Google Document AI, Azure — not yet run), vLLM-based models (Surya2, Unlimited-OCR — planned for a later batch), or fine-tuned research models from the original SROIE competition.

This is a first-party benchmark reference. All figures are traced to frozen benchmark artifacts produced on August 11–12, 2026 under a versioned protocol (see Methodology). The full artifact package — predictions, metrics, manifests, performance logs — is available on request. Where a metric does not apply to a model, it is reported as not applicable rather than zero.

The headline of this benchmark is an inversion: docTR reads receipt text most accurately (19.7% CER) but extracts structured fields the worst (7.7% field F1), while PaddleOCR reads slightly less accurately (20.5% CER) yet extracts fields 4× better (32.5% field F1). Text accuracy is not field accuracy — and any model-selection decision made on a single "accuracy" number misses the split.

19.7%
Lowest Character Error Rate (CER) on SROIE English receipts — docTR v1.0.1, 361 samples, 95% CI 18.5–21.0%
32.5%
Highest field-level F1 (company/date/address/total, regex postprocessed from OCR text) — PaddleOCR 3.7.0, 361 samples
7.7×
Wall-clock speed gap between the fastest (docTR, 256.0 pages/min) and slowest (Docling, 33.2 pages/min) model on identical RTX 4090 hardware — cost per 1,000 pages spans 7.8×

The Paradox: Best Text Reader, Worst Field Extractor

The same benchmark run produced both the best and worst field-extraction result on the same 361 receipts. The difference is not data quality — it is the difference between reading text and extracting fields.

Two accuracy metrics answer different questions. Character Error Rate (CER) measures how many individual characters were misread relative to the ground-truth transcription — a 19.7% CER means roughly one wrong character per five. Word Error Rate (WER) fails an entire word on any single misread character inside it, which is why WER values are always higher than CER. Field-level F1, in this benchmark, measures whether the four SROIE fields — company name, date, address, and total — can be recovered from the OCR text by a fixed regex postprocessor, computed as the harmonic mean of precision and recall across the 361-document field set.

docTR produced the cleanest transcription (19.7% CER, 32.0% WER) but its field recovery collapsed to 7.7% F1 — a regex pass over clean text can still fail when the value format (currency symbols, comma-decimal totals, multi-line addresses) does not match the extraction pattern. PaddleOCR's transcription was slightly noisier (20.5% CER) yet its output aligned better with the field patterns, yielding 32.5% F1. Neither result is a mistake — they are two different layers of the same pipeline, and no model in this batch achieved even one fully correct document on all four fields (document-level field exact match = 0 for every model). This is the practical meaning of the statement that OCR text accuracy is not field accuracy.

SROIE 2019 Results: Text Accuracy, Field F1, Latency, and Cost

The primary ranking set is SROIE 2019 (Scanned Receipt OCR and Information Extraction, the ICDAR 2019 competition dataset) — 361 English-language scanned receipts in the fixed official test split, scored with all five models in the same execution group on the same GPU. Lower CER/WER is better; higher field F1 is better.

Character Error Rate (CER) by model on SROIE 2019 (lower is better): docTR 19.7%, PaddleOCR 20.5%, EasyOCR 28.3%, Tesseract 33.6%, Docling 58.4%.

Source: ImageToTable.ai Receipt OCR Benchmark v1, SROIE 2019 test split (361 samples). Runs: official tier, warm_then_scored protocol, RTX 4090. Bootstrap 95% CI via 2,000 resamples. Full artifact package available on request. Dataset: ICDAR 2019 SROIE competition.

Model (version)CER (95% CI)WER (95% CI)Field F1p50 Latencyp95 LatencyPages/min (wall)$/1K pages
docTR v1.0.119.7% (18.5–21.0)32.0% (30.4–33.5)7.7%153 ms455 ms256.0$0.049
PaddleOCR 3.7.020.5% (19.2–21.7)32.6% (31.0–34.1)32.5%219 ms612 ms157.8$0.080
EasyOCR 1.7.228.3% (27.1–29.5)61.6% (59.5–63.5)14.8%660 ms1,440 ms77.4$0.164
Tesseract 5.3.433.6% (31.2–35.9)56.2% (53.4–58.9)23.2%939 ms2,314 ms54.4$0.233
Docling 2.119.058.4% (52.8–64.6)75.1% (69.6–81.2)22.3%1,196 ms4,834 ms33.2$0.381

Source: ImageToTable.ai Receipt OCR Benchmark v1, SROIE 2019 test split. All metrics from frozen benchmark artifacts (n=361 per model, 100% success rate). CER/WER: character/word error rate against ground-truth transcription. Field F1: postprocessed field extraction from OCR text via fixed regex — NOT native model field output (none of the five models emits native structured fields). CI: bootstrap percentile 95%, 2,000 resamples. Cost: computed from RunPod RTX 4090 at $0.76/hour (price recorded 2026-08-11). Full artifacts available on request.

Why Field F1 Lags Text Accuracy So Far

Every model in this batch reads the text reasonably but fails to deliver a usable structured field on most receipts — the best field F1 is 32.5%, and no model produced a single fully-correct document (company + date + address + total all exact).

The SROIE field metrics here are regex postprocessed: the benchmark takes each model's OCR text and applies a fixed pattern-based extraction for the four fields. This is deliberately different from a model's native structured-field output — none of the five open-source OCR engines in this batch emits native structured fields for receipt layout, so the postprocessor is the only field path available to them. The gap between a 19.7% CER and a 7.7% field F1 is what a real extraction pipeline pays for the missing structured-output layer: high-quality text still needs layout understanding and value normalization to become fields.

The ranking inversion between models is stable: PaddleOCR leads field F1 at 32.5% (30.5–34.5% CI), followed by Tesseract 23.2%, Docling 22.3%, EasyOCR 14.8%, and docTR 7.7% — while the text-metric ranking is almost exactly reversed at the top (docTR 19.7% CER vs PaddleOCR 20.5% CER). Any pipeline that reports only "OCR accuracy" is hiding the field-extraction layer where the real variance lives.

CORD v2: The Language Stress Test

On 100 Indonesian-language receipts (CORD v2), every model's CER collapses to 90–95% — an English-trained OCR stack does not transfer to another language, and this benchmark quantifies the penalty instead of hand-waving it.

CORD v2 (a consolidated receipt dataset built for post-OCR parsing) provides the cross-language stress test. Its 100 reviewed test samples are Indonesian-language receipts, and every model in this batch was trained primarily on English data. The results below are not a ranking of model quality — CORD text metrics are framed as receipt/language/layout robustness evidence and must not be merged with SROIE into a single overall ranking. Field metrics are not applicable on CORD because no model emits native CORD structured fields and no CORD postprocessor is defined for this batch.

Model (version)CERWERp50 Latency$/1K pages
PaddleOCR 3.7.090.8%94.1%123 ms$0.067
docTR v1.0.191.0%93.7%123 ms$0.055
EasyOCR 1.7.291.8%96.2%6,450 ms$1.445
Docling 2.119.092.2%95.3%524 ms$0.244
Tesseract 5.3.495.2%97.7%652 ms$0.161

Source: ImageToTable.ai Receipt OCR Benchmark v1, CORD v2 test split (100 samples). All runs official tier, RTX 4090, 100% success rate. Caveat: language-mismatch stress test — do not compare CORD CER values to SROIE CER values. Dataset: CORD v2 (Clova AI).

Speed and Cost: The 7.7× Spread

On identical hardware, throughput spans 7.7× and cost per 1,000 pages spans 7.8× — for receipt volume, choosing the fastest model can cut OCR compute cost by roughly 87% before accuracy is even considered.

Throughput is reported two ways in this benchmark: per-page p50/p95 latency (from successful prediction records, excluding runner initialization) and end-to-end wall-clock pages per minute (including runner process startup under warm_then_scored mode). Cost per 1,000 pages is computed from run wall time at the recorded RunPod RTX 4090 rate of $0.76/hour. The charts below show SROIE wall-clock throughput and cost; CORD follows the same shape with docTR and PaddleOCR again fastest and cheapest.

Wall-clock throughput on SROIE 2019 (RTX 4090, warm_then_scored): docTR 256.0, PaddleOCR 157.8, EasyOCR 77.4, Tesseract 54.4, Docling 33.2 pages per minute.

Source: ImageToTable.ai Receipt OCR Benchmark v1, SROIE 2019 test split (361 samples). Wall-clock pages per minute from timed benchmark runs including runner initialization. All runs official tier, RTX 4090.

Cost per 1,000 pages on SROIE 2019 (RTX 4090 at $0.76/hr): docTR $0.049, PaddleOCR $0.080, EasyOCR $0.164, Tesseract $0.233, Docling $0.381.

Source: ImageToTable.ai Receipt OCR Benchmark v1, SROIE 2019 test split. Cost = run wall time × $0.76/hr (RunPod RTX 4090, price recorded 2026-08-11). Includes per-run initialization overhead.

How to Pick a Model From These Results

There is no single "best" model in this benchmark — the results support picking by priority, and the same table answers different questions. Three common priorities map directly onto the data:

  • If raw transcription accuracy is the goal (full-text indexing, human-readable digitization): docTR leads at 19.7% CER and is also the fastest and cheapest at 256.0 pages/min and $0.049 per 1,000 pages — it wins all three of text accuracy, speed, and cost simultaneously.
  • If structured field extraction is the goal (feeding downstream systems that need company/date/total): PaddleOCR leads at 32.5% field F1, but even that is a ~2-in-3 field failure rate — the honest conclusion is that none of these open-source OCR engines alone delivers production-grade receipt field extraction without a structured-extraction layer on top.
  • If cost at volume is the constraint: the 7.8× cost spread means a million pages cost roughly $49 with docTR versus $381 with Docling — a difference of ~$332 per million pages on identical hardware.

Whichever priority applies, the protocol and artifacts let you reproduce every number in this table on your own GPU before committing — and for a multilingual receipt flow, the CORD results are the warning that English-trained models should be stress-tested on the target language first.

Frequently Asked Questions

Which open-source OCR model is most accurate on receipts?

It depends on the metric: docTR has the best text accuracy (19.7% CER on 361 SROIE receipts) while PaddleOCR has the best field extraction (32.5% field F1) in the ImageToTable.ai Receipt OCR Benchmark. No single model leads both — text accuracy and field extraction are different pipeline layers.

Why is docTR's text accuracy better but its field extraction worse than PaddleOCR?

Because the two metrics measure different pipeline layers. docTR read characters more accurately (19.7% CER vs 20.5%) but its output did not match the fixed regex field patterns for company/date/address/total, dropping its postprocessed field F1 to 7.7% while PaddleOCR reached 32.5%. Cleaner text does not guarantee better field extraction when the value format (currency symbols, comma separators, multi-line layout) diverges from the extraction pattern.

Is PaddleOCR better than Tesseract for receipt OCR?

Yes on every metric in this benchmark: PaddleOCR beats Tesseract on CER (20.5% vs 33.6%), WER (32.6% vs 56.2%), field F1 (32.5% vs 23.2%), speed (157.8 vs 54.4 pages/min), and cost ($0.080 vs $0.233 per 1,000 pages) on the 361-receipt SROIE test set.

Does OCR accuracy drop for non-English receipts?

Dramatically: on 100 Indonesian CORD v2 receipts, every English-trained model scored 90–95% CER (PaddleOCR 90.8% best, Tesseract 95.2% worst) versus 20–58% on English SROIE receipts. Treat this as a language stress test, not a model ranking — the same models were not trained for Indonesian.

How much does open-source OCR cost per page on a GPU?

Between $0.049 and $0.381 per 1,000 pages on an RTX 4090 at the recorded RunPod rate of $0.76/hour — roughly $0.00005 to $0.0004 per page including run startup, depending on model.

What is SROIE and what does it test?

SROIE 2019 (Scanned Receipt OCR and Information Extraction) is the ICDAR 2019 competition dataset of real scanned English receipts with transcription and key-field ground truth — this benchmark uses its fixed 361-sample test split. It tests whether an OCR model can read receipt text (CER/WER) and recover company, date, address, and total fields; for the open-source models here, fields were recovered via a fixed regex postprocessor because none emits native structured fields.

Can I reproduce these benchmark numbers myself?

Yes. The benchmark protocol is versioned and frozen, all datasets are publicly available (SROIE from ICDAR 2019, CORD v2 from Clova AI), and all five models are open-source with published versions. The full artifact package — predictions, metrics, manifests, performance logs — is available on request. One caveat: current manifests predate the schema v2 environment fingerprint and a rerun in the exact model environments would be needed for full byte-for-byte reproducibility; see Limitations.

Methodology & Sources

Protocol

This benchmark follows a versioned, frozen protocol (v0.1, 2026-08-11). Every number on this page traces to a specific prediction artifact produced under the protocol, and every run satisfies the same set of publication gates: run_tier=official, expected_split=test, no missing predictions, and all not_applicable metrics preserved as-is (not converted to zero). A quality audit (2026-08-12) verified all prediction contracts, sample hashes, and performance hashes.

Runtime Environment

ComponentDetail
GPUNVIDIA GeForce RTX 4090, 24 GB VRAM, driver 570.211.01
GPU providerRunPod, $0.76/hour (price recorded 2026-08-11T18:30:00Z)
Python3.12.3
PyTorch2.8.0+cu128 (EasyOCR, docTR, Docling); not used by Tesseract, PaddleOCR
Measurement modewarm_then_scored: 5-sample warm-up per dataset, followed by full test split
SplitTest-only; no training or validation samples scored
Manual auditAll prediction contracts passed — no missing/dangling/duplicate sample IDs (audit 2026-08-12)

How to Reproduce

The benchmark code, runners, and evaluators are maintained in the ImageToTable.ai benchmark-ocr repository (Git commit 8bdc616; public release pending). The exact commands that produced every number on this page:

SROIE 2019 (English receipts, 361 test samples):

export BENCHMARK_RUN_TIER=official
export BENCHMARK_EXPECTED_SPLIT=test
export BENCHMARK_MEASUREMENT_MODE=warm_then_scored
export BENCHMARK_GPU_LABEL="rtx_4090"
export BENCHMARK_GPU_PROVIDER="runpod"
export BENCHMARK_GPU_HOURLY_USD="0.76"
export BENCHMARK_FIELD_POSTPROCESSOR=sroie_receipt_regex

for model in tesseract paddleocr easyocr doctr docling; do
  BENCHMARK_WARMUP_SAMPLES=sroie_warmup_5.jsonl \
    bash server/run_model.sh "$model" sroie 361 "receipt-v1-sroie-${model}"
done

CORD v2 (Indonesian receipts, 100 test samples):

unset BENCHMARK_FIELD_POSTPROCESSOR

for model in tesseract paddleocr easyocr doctr docling; do
  BENCHMARK_WARMUP_SAMPLES=cord_v2_warmup_5.jsonl \
    bash server/run_model.sh "$model" cord_v2 100 "receipt-v1-cord-v2-${model}"
done

All models used out-of-the-box with default configurations. No fine-tuning, no custom language packs, no post-training adaptations. See the benchmark repository for per-model runner scripts and environment setup.

Metric Definitions

  • CER (Character Error Rate): Levenshtein edit distance at the character level divided by ground-truth character count (lower is better). Bootstrap 95% CI via 2,000 resamples, seed 20260811.
  • WER (Word Error Rate): Levenshtein edit distance at the word level (whitespace tokenization) divided by ground-truth word count (lower is better). Same CI method as CER.
  • Field F1 (SROIE only): harmonic mean of precision and recall across the four SROIE fields (company, date, address, total), extracted from OCR text by a fixed regex postprocessor (sroie_receipt_regex). NOT native model structured-field output — all five models have not applicable for native field metrics.
  • p50 / p95 latency: per-page inference time (ms) from successful prediction records, excluding runner initialization.
  • Wall-clock pages/min: total scored samples divided by total run wall time, including runner process startup. Higher is better.
  • Cost per 1,000 pages: run wall time (hours) × $0.76 × (1,000 / sample count). Uses actual provider price metadata, not estimated rates.
  • Success rate: 100% for all 10 runs in this batch (0 errors, 0 timeouts, 0 missing predictions).

Datasets

  1. SROIE 2019 (Scanned Receipt OCR and Information Extraction) — ICDAR 2019 Robust Reading Competition, Task 3. 361 real scanned English-language receipts with line-level text transcription and four key-field ground-truth labels (company, date, address, total). Fixed test split.
  2. CORD v2 (Consolidated Receipt Dataset) — Clova AI Research, NAVER Corp. 100 reviewed Indonesian-language receipt samples from the public test split, used as a cross-language stress test. Includes nested menu line-item and subtotal/total ground truth (not scored in this batch — requires native structured-field emitters).

Models Tested

ModelVersionTypeSource
Tesseract5.3.4Classic OCR engine (CPU+GPU)GitHub
PaddleOCR3.7.0Deep-learning OCR (PaddlePaddle)GitHub
EasyOCR1.7.2Deep-learning OCR (PyTorch)GitHub
docTRv1.0.1Neural OCR (TensorFlow / PyTorch)GitHub
Docling2.119.0Document parser (IBM)GitHub

Artifact Access & Reproducibility

Benchmark source code: maintained in the benchmark-ocr repository (Git commit 8bdc616). The repository includes per-model runner scripts, evaluators (CER/WER, field metrics, bootstrap CI), dataset sample manifests with SHA256 hashes, and the frozen protocol. Public release is pending — once published, every number on this page will be independently reproducible by cloning the repository, obtaining the public datasets, and running the commands above on an RTX 4090.

Full artifact package: per-model predictions (*.jsonl), metrics with 95% CI (metrics.csv), field-level diagnostics (field_details.csv), wall-clock performance logs (performance.json), and run manifests (manifest.json) are versioned under the protocol and available alongside the source code. Contact [email protected] for early access before public release.

Schema v1 note: current manifests predate the schema v2 environment fingerprint (which captures OS, CUDA version, and per-model Python package versions at runtime). The environment table above reflects what schema v1 records; full byte-for-byte reproducibility will require a schema v2 rerun. This does not affect the correctness of the reported metrics.

Limitations

  • Manifest schema v1: current run manifests predate the schema v2 environment fingerprint, which captures the exact Python environment at runtime. Full byte-for-byte reproducibility requires a schema v2 rerun in the exact model environments.
  • SROIE field metrics are regex postprocessed, not native extraction: field F1 measures field recovery from OCR text via fixed patterns — it is not the models' native structured-field capability (which is not available for any model in this batch).
  • CORD results are a language-mismatch stress test, not an accuracy ranking: all five models are English-trained; CORD CER values (90%+) quantify cross-language collapse and must not be merged with SROIE into a single leaderboard.
  • PaddleOCR-VL blocked from ranking: excluded due to direct pipeline issues; its results are not comparable to the five ranked models.
  • Surya2 and Unlimited-OCR not included: they require a vLLM server Pod and are planned for a v1.1 batch; the current release covers the local_torch group only.
  • Single GPU tier only: all runs are RTX 4090; no multi-GPU or RTX 3090/A100 comparison is available yet, so cost/throughput conclusions are specific to this one tier.
  • Open-source models only: no cloud/API models (AWS Textract, Google Document AI, Azure Form Recognizer) are included; their comparison is a separate planned benchmark.

Related references: Receipt OCR Accuracy · OCR Accuracy by Document Type · What is OCR?

Related reading: ABBYY FineReader vs Modern AI OCR · Adobe Acrobat OCR vs AI Extraction · How to Read OCR Accuracy Claims

📮 contact email: [email protected]