Tesseract vs PaddleOCR on ReceiptsLegacy CPU vs Modern GPU (2026)

Last reviewed: 2026-08-18 · Run tier: official · First-party head-to-head benchmark · 2 engines × 2 receipt datasets

What this page covers: A first-party, reproducible head-to-head between Tesseract 5.3.4 (classic open-source OCR, ~35 years of lineage, CPU-only in this benchmark) and PaddleOCR 3.7.0 (modern two-stage deep-learning OCR — PP-OCR detection + recognition — running on GPU), on two receipt datasets: SROIE 2019 English receipts (361 test samples) and CORD v2 Indonesian receipts (100 test samples). Metrics compared per engine: character error rate (CER), word error rate (WER), field-extraction F1 under two postprocessing methods (fixed regex patterns and an LLM), p50/p95 latency, wall-clock pages per minute, and cost per 1,000 pages — with Tesseract’s CPU-only compute type explicitly separated from GPU-billed numbers throughout. Every number traces to a published CSV row in the public OCR benchmark repository (ImageToTableai/benchmark-ocr) — reproducible experimental data, not an aggregation of third-party reports.
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. Cloud/API OCR services, fine-tuned engines, other open-source engines (only these two are compared), and any hardware tier other than the single recorded RTX 4090 are out of scope except where quoted as ranking context. The complete 8-engine roundup lives on traditional OCR and VLM parsing head to head.

Range statement: every number on this page applies to receipts only — SROIE 2019 English receipts and CORD v2 Indonesian receipts. One hardware tier (RTX 4090 at $0.76/hr, price timestamped August 2026), one LLM postprocessor (deepseek-v4-flash at temperature 0), fixed model versions (Tesseract 5.3.4, PaddleOCR 3.7.0). Tesseract ran on CPU against GPU-accelerated engines — that asymmetry is inherent to the comparison, not a flaw in it. Do not extrapolate these results to other document types, GPUs, or LLMs. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.

The accuracy gap between the generations is decisive and one-sided — not a tie like the earlier docTR-vs-Surya2 head-to-head. On the same 361 SROIE receipts, PaddleOCR wins every accuracy axis: CER 0.2045 vs 0.3347 (39% lower), WER 0.3256 vs 0.5591 (42% lower), regex field F1 0.3254 vs 0.2335 (1.39×), and LLM postprocessed field F1 0.5810 vs 0.4389 (1.32×). But the headline surprise is the opposite direction: a CPU-only classic engine matches the modern GPU engine on wall-clock throughput — 78.6 vs 79.7 pages/min — and holds a 2.2× tighter p95 tail (1,507.0 vs 3,331.4 ms), while costing nothing in GPU billing where PaddleOCR charges $0.2214 per 1,000 pages. The modern engine is not “faster at volume” — it is faster per page once warm, and that advantage is what the wall clock partly eats.

The trade, in one pair of numbers: PaddleOCR reads a receipt with 39% fewer character errors and extracts 1.39× as many fields through regex for $0.2214 per 1,000 pages; Tesseract reads it on CPU with more errors, zero GPU billing (its cost cell is empty by design), and statistically identical wall-clock throughput. Neither engine “wins”; they win on different axes — and on the field-extraction axis the gap widens to the largest sibling-row split in the whole eight-engine benchmark (CORD LLM field F1 0.5527 vs 0.1627).

0.2045 · 0.3347
SROIE CER for PaddleOCR vs Tesseract — a 39% relative gap, the modern two-stage architecture’s text-accuracy edge on English receipts; PaddleOCR ranks 3rd of 8 engines on this metric, Tesseract mid-pack at 5th (summary_metrics.csv, cer, paddleocr/sroie_2019 and tesseract/sroie_2019 rows)
78.6 · 79.7 pg/min
Wall-clock throughput on SROIE — statistical parity between the CPU-only classic engine and the GPU engine (~1.4% apart), while Tesseract also holds a 2.2× tighter p95 tail (summary_metrics.csv, pages_per_minute / latency_p95_ms, same two rows)
CPU-only · $0.2214
Cost per 1,000 pages — Tesseract’s cost cell is empty (no GPU billing, CPU-only by design, not zero); PaddleOCR bills $0.2214 on the same RTX 4090 at $0.76/hr (summary_metrics.csv, cost_per_1000_pages, same two rows)

What the Two Engines Are: 35 Years of OCR vs a CNN Two-Stage Pipeline

The entire story of this page is an architecture gap. Tesseract is the classic open-source OCR engine — originally developed at HP in the 1980s and open-sourced by Google in 2005, which is why it carries roughly 35 years of lineage. Its pipeline is traditional computer vision: adaptive binarization, page segmentation, connected-component analysis, and character recognition — LSTM-based since version 4 — all executing on CPU with no GPU billing in this benchmark (compute_type = cpu in the CSV). PaddleOCR is a modern deep-learning engine from the PaddlePaddle ecosystem: a two-stage pipeline in the PP-OCR family — a detection stage that localizes text regions (DBNet-style), then a recognition stage that transcribes them — running on GPU. One engine reads by matching character shapes against learned patterns; the other reads by learning where text is and what it says. This benchmark puts both through the same receipts, same protocol, same machine.

Why this mechanism matters: Tesseract’s approach is cheap to run and needs no GPU — but its character model is frozen decades of classic recognition, which shows up as a hard ceiling on text quality. PaddleOCR’s approach costs GPU time but reads substantially cleaner text. The benchmark’s job is to put a number on both sides of that trade from one controlled run — and the surprise is how narrow the operating-cost side of the trade turned out to be.

Character Accuracy: The Modern Architecture Wins Every Text Metric

On SROIE 2019, the text-accuracy gap is large and one-sided: CER 0.2045 (PaddleOCR) vs 0.3347 (Tesseract) — a 39% relative improvement — and WER 0.3256 vs 0.5591, a 42% relative gap. Tesseract’s CER of 0.3347 ranks fifth of the eight engines in the underlying run — mid-pack, not last — but every engine above it except one is a deep-learning engine, and the gap between Tesseract and the deep-learning tier (best: Surya2 0.1915, docTR 0.1971) is larger than the gap between those engines and PaddleOCR (0.2045, third).

Character Error Rate (CER) is the classic OCR yardstick: insertions, deletions, and substitutions divided by ground-truth characters — a CER of 0.335 means roughly 33.5 misread characters per 100. Word Error Rate (WER) applies the same edit-distance calculation at word granularity. Both are lower-is-better. The WER gap (42%) being wider than the CER gap (39%) means Tesseract’s character slips compound into whole-word failures on this corpus — the classic-engine failure mode that a downstream field extractor inherits directly.

Text accuracy on SROIE 2019: PaddleOCR CER 20.4% vs Tesseract 33.5%; WER 32.6% vs 55.9%. Lower is better. A 39% relative CER gap and a 42% WER gap.

Source: summary_metrics.csv — cer and wer columns, sroie_2019 rows. PaddleOCR cer 0.20449 / wer 0.32563; Tesseract cer 0.33468 / wer 0.55915. Lower is better. 361 samples per engine; both error_rate 0.0.

Metric (SROIE 2019, n=361)Tesseract 5.3.4 (CPU)PaddleOCR 3.7.0 (GPU)Source
Character Error Rate (CER)0.33470.2045summary_metrics.csv · cer, tesseract/sroie_2019 and paddleocr/sroie_2019 rows
Word Error Rate (WER)0.55910.3256summary_metrics.csv · wer, same rows
Error rate (failed pages)0.00.0summary_metrics.csv · error_rate, same rows

Table: summary_metrics.csv — cer / wer / error_rate columns, sroie_2019 rows. Exact values: Tesseract cer 0.33468 / wer 0.55915; PaddleOCR cer 0.20449 / wer 0.32563. Lower CER/WER is better. Ranking context from the same CSV: SROIE CER across all eight engines runs surya2 0.1915, doctr 0.1971, PaddleOCR 0.2045 (3rd), easyocr 0.2833, Tesseract 0.3347 (5th), paddleocr_vl 0.3370, docling 0.5909, unlimited_ocr 0.6552. Neither engine here is the benchmark’s accuracy champion — docTR and Surya2 hold the top two CER slots.

Field Extraction: The Gap That Decides Production Decisions

Text accuracy ranks engines; field extraction is what downstream systems actually consume. The benchmark’s SROIE field metrics target four flat receipt fields (company, date, address, total) using two postprocessors on each engine’s OCR text: fixed regex patterns (the traditional OCR + rule-based key-information-extraction approach) and an LLM postprocessor (deepseek-v4-flash at temperature 0) with a structured prompt. Through regex, PaddleOCR extracts fields at 0.3254 field F1 to Tesseract’s 0.2335 — a 1.39× advantage; through the LLM, the gap persists at 0.5810 vs 0.4389 (1.32×). The LLM lever helps both engines, but it starts from Tesseract’s weaker base — and the ceiling argument that decides production pipelines lives one section below.

Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth — 1.0 means every receipt field perfectly recovered, 0 means nothing. The SROIE regex field columns are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text — postprocessed, not native structured output. Ranking context: PaddleOCR’s regex F1 of 0.3254 is the best among the four purely traditional engines in the 8-engine benchmark (behind only Unlimited-OCR 0.3376 and PaddleOCR-VL 0.3368); Tesseract’s 0.2335 is the fourth-best regex result overall (summary_metrics.csv, field_f1_regex, sroie_2019 rows) — the classic engine is a middle-of-the-pack field extractor on clean English text, which is precisely where it stops being competitive.

SROIE 2019 field F1 by postprocessing method: through regex patterns PaddleOCR reaches 32.5% vs Tesseract 23.3%; through LLM postprocessing (deepseek-v4-flash) PaddleOCR 58.1% vs Tesseract 43.9%.

Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).

Field extraction (SROIE 2019, n=361)Tesseract 5.3.4 (CPU)PaddleOCR 3.7.0 (GPU)Source
Field-value F1 (regex)0.23350.3254field_method_comparison.csv · regex_field_value_f1, tesseract/sroie_2019 and paddleocr/sroie_2019 rows
Field-value F1 (LLM)0.43890.5810field_method_comparison.csv · llm_field_value_f1, same rows
Docs with all fields exact (LLM)0.05260.0748field_method_comparison.csv · llm_document_fields_exact, same rows
Median LLM postprocess latency (ms)1,837.11,817.5field_method_comparison.csv · llm_median_latency_ms, same rows

Table: field_method_comparison.csv — regex and llm columns, sroie_2019 rows. The regex columns are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text. LLM postprocessor: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms). “Docs with all fields exact” is the fraction of documents where every target field matched exactly — a much harsher bar than per-field F1. Ranking context (llm_field_value_f1, all sroie_2019 rows): PaddleOCR 0.5810 ranks 5th of 8; Tesseract 0.4389 ranks 7th, ahead of only EasyOCR’s 0.3717.

The Surprise: CPU Throughput Parity on the Wall Clock

The page’s headline-worthy finding is the one nobody’s third-party comparison documents: on the same receipts, a CPU-only classic engine matches a modern GPU engine on wall-clock pages per minute — 78.6 (Tesseract) vs 79.7 (PaddleOCR), about 1.4% apart, statistical parity. The classic engine is not “slow at volume”: it is slow per page but steady — and on this corpus it outruns four of the seven GPU engines in the underlying run (docling 56.7, paddleocr_vl 68.2, unlimited_ocr 34.4, surya2 12.1 pages/min).

This looks like a contradiction with the latency numbers, and it deserves an honest reconciliation rather than a footnote. Latency p50 is steady-state per-page inference, measured warm-then-scored with model loading excluded — PaddleOCR’s 297.0 ms is genuinely faster than Tesseract’s 670.9 ms. Pages per minute is wall-clock throughput of the whole run, including model initialization and batch effects. Convert the CSV’s throughput to wall-clock time per page (60 seconds ÷ pages_per_minute): PaddleOCR spends ~753 ms per page wall-clock against a 297 ms p50 — about 456 ms per page of init/prefill and batch overhead; Tesseract spends ~763 ms per page wall-clock against a 671 ms p50 — about 92 ms of overhead. Tesseract’s stripped CPU runtime starts fast and streams steadily; PaddleOCR’s GPU pipeline pays a heavier load/prefill price per run that nearly cancels its fast steady-state on this 361-page corpus. A long-running warm pipeline sees PaddleOCR’s per-page edge; a pipeline dominated by cold starts, small batches, or frequent re-initialization sees the two engines at parity or better on the classic side.

The p95 tail tells the same story in one number: Tesseract’s p95 of 1,507.0 ms is 2.2× tighter than PaddleOCR’s 3,331.4 ms. The GPU engine’s first-page/prefill spike — the same load path that inflates its wall-clock per-page time — dominates its worst-case tail, while the CPU engine has no such spike. For tail-latency-sensitive or capacity-planned workloads, the classic engine is the more predictable one.

Wall-clock throughput on SROIE 2019: Tesseract 78.6 pages/min vs PaddleOCR 79.7 pages/min — about 1.4% apart, statistical parity between a CPU-only classic engine and a GPU engine.

Source: summary_metrics.csv — pages_per_minute column, sroie_2019 rows. Tesseract 78.63285, PaddleOCR 79.71298. Wall-clock pages/min including model initialization; per-page steady-state latency is the latency_p50_ms column (see chart below). Reconciliation: 60 ÷ 78.63285 = 763 ms/page vs 60 ÷ 79.71298 = 753 ms/page wall-clock.

Latency on SROIE 2019: PaddleOCR p50 297.0 ms (2.3x faster per page once warm) but p95 3,331.4 ms (2.2x wider tail); Tesseract p50 670.9 ms but p95 1,507.0 ms — tighter tail, no first-page prefill spike. Steady-state, warm-then-scored (excludes model loading).

Source: summary_metrics.csv — latency_p50_ms / latency_p95_ms columns, sroie_2019 rows. Tesseract p50 670.87 / p95 1506.99; PaddleOCR p50 296.99 / p95 3331.35. Steady-state latency (warm_then_scored measurement mode, excludes model loading). The p50-vs-pages/min tension is reconciled in the prose above: different clocks, both real.

Operating envelope (SROIE 2019, n=361)Tesseract 5.3.4 (CPU)PaddleOCR 3.7.0 (GPU)Source
Latency p50 (ms)670.9297.0summary_metrics.csv · latency_p50_ms, tesseract/sroie_2019 and paddleocr/sroie_2019 rows
Latency p95 (ms)1,507.03,331.4summary_metrics.csv · latency_p95_ms, same rows
Pages per minute (wall-clock)78.679.7summary_metrics.csv · pages_per_minute, same rows

Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute, sroie_2019 rows. Exact values: Tesseract p50 670.87 / p95 1506.99 / 78.63 pg/min; PaddleOCR p50 296.99 / p95 3331.35 / 79.71 pg/min. Latency is steady-state per-page (warm-then-scored, excludes model loading); pages/min is wall-clock including init and batch effects — the parity and the p95 inversion are measurement-model and architecture facts, not contradictions.

The Cost Story: Where the Legacy Engine Wins Outright

Cost is the one axis where Tesseract’s age is an advantage, and it is a structural one: Tesseract is CPU-only, so its cost cell in the CSV is empty by design — no GPU billing to meter — while PaddleOCR bills $0.2214 per 1,000 pages on the same RTX 4090 at the recorded $0.76/hr. For a throughput-tied workload (the parity above), the classic engine’s operating cost on GPU-billing-shy infrastructure is materially lower — the price anchor for the “is upgrading worth it” decision.

Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization. The empty Tesseract cell is not a zero — it is a missing value because the engine never touched the GPU; the benchmark records it as empty rather than assuming a number (protocol rule: an empty cell is not_applicable, never 0). Two context numbers keep this honest: PaddleOCR’s $0.2214 is mid-pack among the seven GPU engines (docTR holds the benchmark’s cheapest GPU row at $0.048 per 1,000 pages), and on CORD PaddleOCR’s cost rises to $0.3419 per 1,000 pages at 141.0 pages/min.

Cost per 1,000 pages on SROIE 2019 (RTX 4090 at $0.76/hr): PaddleOCR $0.2214; Tesseract bar omitted — CPU-only, no GPU cost (cost cell empty in the CSV, not zero).

Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. PaddleOCR 0.2214. Tesseract’s value is empty (cell blank in the CSV): CPU-only compute type, no GPU billing — plotted as omitted, not zero. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026). Cheapest GPU engine in the benchmark: docTR $0.048 per 1,000 pages (doctr/sroie_2019 row).

Cost & throughput (SROIE 2019, n=361)Tesseract 5.3.4 (CPU)PaddleOCR 3.7.0 (GPU)Source
Cost per 1,000 pagesempty — CPU-only (no GPU cost)$0.2214summary_metrics.csv · cost_per_1000_pages, same rows; Tesseract cell blank by design
Compute typecpugpusummary_metrics.csv · compute_type, same rows

Table: summary_metrics.csv — cost_per_1000_pages / compute_type columns, sroie_2019 rows. Tesseract’s cost cell is empty (blank, not 0.0000) because the engine is CPU-only; PaddleOCR’s GPU cost includes model initialization at the recorded $0.76/hr. On CORD, PaddleOCR’s cost is $0.3419 per 1,000 pages at 141.0 pages/min (paddleocr/cord_v2 row).

CORD (Indonesian Receipts): Both Collapse — and the Widest Field Gap in the Benchmark Opens

Neither engine was trained predominantly on Indonesian receipts, so CORD v2 (100 samples, nested fields menu/sub_total/total) functions as a cross-language stress test — and both collapse on raw CER: 0.9083 (PaddleOCR) and 0.9523 (Tesseract), a language-mismatch wash. Per the benchmark protocol, CORD numbers are kept quarantined from the SROIE comparison — never merged into any ranking — because CORD’s ground-truth text embeds annotation structure, which inflates raw CER for every engine on top of the genuine language mismatch.

Where the two engines really separate is the LLM field lever, and this is the page’s strongest single data point: through the LLM postprocessor, PaddleOCR’s CORD field F1 holds at 0.5527 — the best of all eight engines on CORD — while Tesseract’s collapses to 0.1627, the worst of all eight engines in the entire benchmark. That 0.39-point spread is the widest gap between any two sibling LLM-field rows in the run. The mechanism is the ceiling argument made concrete: Tesseract’s CORD text is unreadable enough (CER 0.9523) that no postprocessor — regex or LLM — can recover fields from it. The LLM lever helps (Tesseract’s SROIE F1 rises from 0.2335 regex to 0.4389 LLM), but it starts from a weaker base and cannot manufacture text the engine never read. CORD is quoted here for language-robustness context; it is deliberately never pooled with the SROIE numbers into a single leaderboard.

CORD v2, Indonesian receipts (n=100)Tesseract 5.3.4 (CPU)PaddleOCR 3.7.0 (GPU)Source
Character Error Rate (CER)0.95230.9083summary_metrics.csv · cer, tesseract/cord_v2 and paddleocr/cord_v2 rows
Field-value F1 (regex)0.07520.0154field_method_comparison.csv · regex_field_value_f1, same rows
Field-value F1 (LLM)0.16270.5527field_method_comparison.csv · llm_field_value_f1, same rows
Cost per 1,000 pagesempty — CPU-only$0.3419summary_metrics.csv · cost_per_1000_pages, same rows
Pages per minute (wall-clock)108.9141.0summary_metrics.csv · pages_per_minute, same rows

Table: summary_metrics.csv (cer / cost_per_1000_pages / pages_per_minute) and field_method_comparison.csv (field F1), cord_v2 rows. Do not merge CORD numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth, and the regex patterns were written for English formats (both engines’ regex F1 collapses to ~1–8%). Ranking context (llm_field_value_f1, all cord_v2 rows): PaddleOCR 0.5527 is the best of eight; Tesseract 0.1627 is the worst of eight — the widest sibling-row gap in the benchmark. Tesseract’s cost cell is empty (CPU-only), never 0.

Who Wins When: The Recap Grid

“Better” is workload-dependent, and this head-to-head splits the axes with unusual clarity: every accuracy axis favors PaddleOCR; cost, CPU-simplicity, and the tight p95 tail favor Tesseract; wall-clock throughput is a statistical tie; per-page latency favors PaddleOCR; and the raw-CER championship belongs to neither (docTR/Surya2).

Text accuracy — PaddleOCR
CER 0.2045 vs 0.3347
SROIE character error rate, 39% relative gap; WER 0.3256 vs 0.5591, 42% relative gap (summary_metrics.csv, cer / wer, sroie_2019 rows). PaddleOCR ranks 3rd of 8 on SROIE CER; Tesseract mid-pack at 5th.
Field extraction — PaddleOCR
1.39× regex F1 · 1.32× LLM F1
SROIE field F1 through regex 0.3254 vs 0.2335 and through the LLM 0.5810 vs 0.4389 (field_method_comparison.csv, regex_field_value_f1 / llm_field_value_f1, sroie_2019 rows). PaddleOCR wins the extraction pipeline under both postprocessors.
LLM downstream — PaddleOCR
0.5527 vs 0.1627 F1
CORD LLM field F1: PaddleOCR is the best of all eight engines, Tesseract the worst of all eight — the widest sibling-row gap in the benchmark, proof that no postprocessor fixes unreadable text (field_method_comparison.csv, llm_field_value_f1, cord_v2 rows).
Per-page latency — PaddleOCR
297.0 vs 670.9 ms
SROIE p50 steady-state latency, 2.3× faster once warm (summary_metrics.csv, latency_p50_ms, sroie_2019 rows). For an interactive per-page wait: 0.3 s vs 0.7 s.
Wall-clock throughput — Tie
79.7 vs 78.6 pg/min
SROIE wall-clock pages per minute — about 1.4% apart, statistical parity between a CPU-only classic engine and a GPU engine (summary_metrics.csv, pages_per_minute, sroie_2019 rows). The p50-vs-pages/min tension is reconciled in the Throughput Parity section above: different clocks, both real.
Tail latency — Tesseract
1,507.0 vs 3,331.4 ms
SROIE p95 — Tesseract’s tail is 2.2× tighter; the GPU engine’s first-page/prefill spike dominates its worst case (summary_metrics.csv, latency_p95_ms, sroie_2019 rows). For capacity-planned workloads, the classic engine is the more predictable one.
Cost & CPU simplicity — Tesseract
CPU-only · vs $0.2214
Tesseract’s cost cell is empty by design — CPU-only, no GPU billing; PaddleOCR bills $0.2214 per 1,000 pages on the same RTX 4090 at $0.76/hr (summary_metrics.csv, cost_per_1000_pages, sroie_2019 rows; Tesseract cell blank, not zero).
Raw-CER championship — Neither
0.1915 · 0.1971
The benchmark’s best character readers are Surya2 (CER 0.1915) and docTR (0.1971) in the same 8-engine run; PaddleOCR (0.2045) is third, Tesseract (0.3347) fifth (summary_metrics.csv, cer, sroie_2019 rows). This page compares the “classic default vs modern default” trade, not the accuracy championship — and the cheapest GPU engine is docTR at $0.048 per 1,000 pages.

Frequently Asked Questions

Is PaddleOCR more accurate than Tesseract on receipts?

Yes — on every accuracy axis measured in this benchmark. On SROIE 2019: CER 0.2045 vs 0.3347 (39% relative improvement), WER 0.3256 vs 0.5591 (42%), regex field F1 0.3254 vs 0.2335 (1.39×), LLM postprocessed field F1 0.5810 vs 0.4389 (1.32×) (summary_metrics.csv and field_method_comparison.csv, sroie_2019 rows). Neither engine is the benchmark’s overall text champion — Surya2 (CER 0.1915) and docTR (0.1971) hold that title.

Why does Tesseract match PaddleOCR on pages per minute despite being CPU-only?

Because pages per minute is wall-clock throughput, not per-page inference speed. PaddleOCR’s steady-state p50 (297.0 ms) is genuinely 2.3× faster than Tesseract’s (670.9 ms), but wall-clock throughput includes model initialization and batch effects: at 79.7 pages/min PaddleOCR spends ~753 ms per page wall-clock against its 297 ms p50, while Tesseract’s stripped CPU runtime spends ~763 ms per page against its 671 ms p50 — the GPU engine pays a heavier load/prefill price per run that nearly cancels its speed advantage on this 361-page corpus (summary_metrics.csv, pages_per_minute / latency_p50_ms, sroie_2019 rows).

Why is Tesseract’s p95 latency tighter than PaddleOCR’s?

Because the GPU engine’s first-page/prefill spike dominates its worst-case tail: PaddleOCR p95 3,331.4 ms vs Tesseract p95 1,507.0 ms, a 2.2× inversion of the p50 ordering (summary_metrics.csv, latency_p95_ms, sroie_2019 rows). Tesseract’s CPU pipeline has no load spike and streams steadily; PaddleOCR’s fast steady-state comes with a heavier initialization path on every run. The two clocks measure different things and both are real.

Is Tesseract cheaper than PaddleOCR?

On GPU billing, yes — Tesseract has none: it is CPU-only, so its cost cell in the CSV is empty by design (never 0), while PaddleOCR bills $0.2214 per 1,000 pages on SROIE and $0.3419 on CORD, on the same RTX 4090 at $0.76/hr with cost including model initialization (summary_metrics.csv, cost_per_1000_pages, sroie_2019 and cord_v2 rows). The benchmark’s cheapest GPU engine overall is docTR at $0.048 per 1,000 pages.

Why is Tesseract’s CORD LLM field F1 the worst in the whole benchmark?

Because the LLM postprocessor cannot recover text the OCR engine never read. Tesseract’s CORD CER is 0.9523 — effectively unreadable on Indonesian receipts — so its LLM field F1 collapses to 0.1627, the worst of all eight engines, while PaddleOCR’s holds at 0.5527, the best of eight (field_method_comparison.csv, llm_field_value_f1, cord_v2 rows). The same lever on clean English text (SROIE) lifts Tesseract to 0.4389 — but the ceiling is set by the base text quality.

Why do both engines score so badly on CORD receipts?

Two compounding causes that the protocol keeps apart from the SROIE ranking: a genuine language mismatch (Indonesian receipts outside both engines’ training focus) and annotation-structure inflation inside CORD’s ground-truth text — CER lands at 0.9083 (PaddleOCR) and 0.9523 (Tesseract) (summary_metrics.csv, cer, cord_v2 rows). What still separates them is LLM-downstream recovery: PaddleOCR 0.5527 vs Tesseract 0.1627 field F1 — the widest sibling-row gap in the benchmark. CORD rows are quoted with framing and never pooled into any combined ranking.

Which engine should a receipt pipeline pick, Tesseract or PaddleOCR?

If your pipeline consumes fields — extracted values for company, date, totals — PaddleOCR is the clear default on receipts: 1.39× field F1 under regex, 1.32× under an LLM, and LLM-downstream recovery that survives the CORD language shock (field_method_comparison.csv). If you need high-volume raw text on clean English documents with zero GPU cost, CPU-only infrastructure, or tail-latency predictability, Tesseract remains a legitimate option: wall-clock throughput ties (79.7 vs 78.6 pages/min), p95 is 2.2× tighter, and there is no GPU billing — but budget for a mid-pack text base (CER 0.3347, 5th of 8) that caps every downstream field pipeline. These results hold for English and Indonesian receipts on one GPU tier in August 2026; re-run on your target corpus before production decisions (see Limitations).

Where do the numbers on this page come from?

Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, regex field F1, latency, cost, throughput; tesseract rows carry compute_type=cpu and an empty cost cell) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.

Methodology & Sources

Protocol

This page reports a head-to-head slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Both engines saw the same images, same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). Both runs completed with error_rate 0.0 on both datasets (summary_metrics.csv error_rate column). The CPU/GPU asymmetry is inherent to this comparison: Tesseract ran on CPU (compute_type=cpu) against GPU-accelerated engines, by design — its cost cell is empty because no GPU time was billed, and its latency/throughput were measured on the same machine under the same protocol. The underlying run contains eight engines total; this page compares only the two named engines, with other engines quoted solely as ranking context. The complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.

Runtime Environment

  • Hardware: both engines ran on the same machine with an NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026). Tesseract ran on CPU and incurred no GPU cost; its cost cell is empty by design.
  • Engines: out-of-the-box, no fine-tuning. Versions locked: Tesseract 5.3.4 (classic open-source OCR engine — traditional CV pipeline with LSTM-based recognition, CPU-only, system python3 runner, no GPU env) and PaddleOCR 3.7.0 (modern two-stage deep-learning OCR — PP-OCR detection + recognition, GPU) — per the public repo model table (README.md) and run manifests.
  • LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on both engines.
  • Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost; not applicable to Tesseract (CPU-only).
  • Field postprocessing: SROIE regex field metrics are postprocessed_sroie_receipt_regex_* (field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set. They measure OCR + downstream extraction, not native structured output by either model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended.

Metric Definitions

  • CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better.
  • WER (Word Error Rate): the same edit-distance calculation at word granularity.
  • Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
  • Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
  • Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
  • Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init. They measure different clocks; the p50-vs-pages/min parity on this page is a measurement-model fact, not an error, and is reconciled in the Throughput Parity section.
  • Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate, including model initialization. Tesseract’s cell is empty (CPU-only) — an empty cell is not_applicable, never 0.

Source List

  1. summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the tesseract and paddleocr rows here (tesseract: compute_type=cpu, cost cell empty).
  2. field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the tesseract and paddleocr rows here (and to all eight sroie_2019 / cord_v2 rows in the ranking context).
  3. ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
  4. results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
  5. Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
  6. Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).

Limitations

  • CPU/GPU asymmetry is inherent, not a flaw: Tesseract ran on CPU against GPU-accelerated engines. Its cost cell is empty by design (no GPU billing, never 0), and its latency/throughput are CPU numbers measured on the same machine under the same protocol — a different machine class could shift its operating envelope. Treat the cost comparison as “GPU billing vs none,” not as a hardware-independent truth.
  • Document scope — receipts only: SROIE + CORD. Nothing here measures Tesseract’s behavior on complex layouts, tables, handwriting, or long documents — document types where classic engines are known to degrade further — nor PaddleOCR’s PP-Structure layout/table capabilities. Do not use this page to conclude either engine “wins on everything.”
  • Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths should be treated as noise, not engineering truth — though the gaps documented here (39% CER, 1.39× regex F1, the 0.39-point CORD LLM-F1 spread) are far beyond that band.
  • Single GPU tier and single price: all GPU numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
  • Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the 1.32× SROIE and 0.39-point CORD gaps may move at the margins. LLM latency (~1,817–1,837 ms median on SROIE, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of either engine’s own latency.
  • Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes.
  • CORD CER is not a per-model quality reading: CORD ground truth embeds annotation structure and neither engine was trained predominantly on Indonesian; CORD CER (0.91–0.95) reflects language mismatch + ground-truth inflation. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
  • Version pinning: results hold for Tesseract 5.3.4 and PaddleOCR 3.7.0 (August 2026). Newer releases of either engine may shift every number on this page; Tesseract’s results specifically reflect 5.3.4 and were re-verified in a 2026-08-14 repeat run.

Related references: PaddleOCR vs EasyOCR Receipt Benchmark · docTR vs Surya2 Receipt Benchmark · the eight-engine comparison of OCR and VLMs · fixed rules vs language models for fields · why a correct character count can still mean a wrong field

Related reading: the accuracy gap between AI and traditional OCR · why AI extraction beats OCR on images · AI Document Extraction Pricing (2026)

📮 contact email: [email protected]