PaddleOCR vs EasyOCR on Receipts
Accuracy vs Cost Benchmark (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party head-to-head benchmark · 2 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. Cloud/API OCR services, fine-tuned engines, and the benchmark’s other six engines (Tesseract, docTR, Docling, Surya2, Unlimited-OCR, PaddleOCR-VL) are out of scope except where quoted as ranking context. The complete 8-engine roundup lives on the differences between traditional OCR and VLM parsing.
Range statement: every number on this page applies to receipts only — SROIE 2019 English receipts and CORD v2 Indonesian receipts. One hardware tier (RTX 4090 at $0.76/hr, price timestamped August 2026), one LLM postprocessor (deepseek-v4-flash at temperature 0), fixed model versions (PaddleOCR 3.7.0, EasyOCR 1.7.2). Do not extrapolate these results to other document types, GPUs, or LLMs — the benchmark measures receipt OCR and receipt-field extraction only. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.
On clean English receipts, the modern two-stage architecture’s advantage is unambiguous — not a tie like the earlier docTR-vs-Surya2 head-to-head. PaddleOCR beats EasyOCR on every accuracy axis on SROIE 2019: CER 0.2045 vs 0.2833 (27.8% relative improvement), WER 0.3256 vs 0.6158 (1.9×), regex field F1 0.3254 vs 0.1477 (2.2×), and LLM postprocessed field F1 0.5810 vs 0.3717 (1.56×). But the trade does not end there: EasyOCR is ~2× cheaper per 1,000 pages ($0.110 vs $0.221), faster on wall-clock throughput (124.5 vs 79.7 pages/min), and tighter in the p95 tail (960.4 vs 3,331.4 ms) — while its text, fed to the same LLM postprocessor, extracts fields worse than any of the benchmark’s eight engines, a paradox this page documents with the raw rows.
The trade, in one pair of numbers: PaddleOCR reads a receipt with 28% fewer character errors and extracts 2.2× as many fields through regex for $0.221 per 1,000 pages; EasyOCR reads it with more errors but for $0.110 per 1,000 pages — roughly half the GPU cost on the same RTX 4090, same test split, same receipts. Neither engine “wins”; they win on different axes, and the point of this page is to show both axes from the same controlled run — including the counter-intuitive LLM-downstream result.
The two engines represent two generations of deep-learning OCR. PaddleOCR (PaddlePaddle-based, PP-OCR architecture, v3.7.0) is a modern two-stage pipeline: a detection stage localizes text regions, then a recognition stage transcribes them — engineered for strong accuracy on printed text at scale. EasyOCR (PyTorch-based, 1.7.2) is a classic single-pass CNN + RNN + CTC recognizer — a ResNet feature extractor feeding a sequence model decoded with Connectionist Temporal Classification. It is known for very wide language and script coverage (80+ languages out of the box) and a famously simple install. Character Error Rate (CER) measures insertions, deletions, and substitutions divided by ground-truth characters — a CER of 0.204 means ~20.4 characters misread per 100; Word Error Rate (WER) applies the same edit-distance logic at whole-word granularity. Lower is better on both.
Text Accuracy on SROIE (English Receipts): PaddleOCR’s Clear Edge
On the 361 English receipts of the SROIE 2019 test split, the modern architecture wins on both text metrics: CER 0.2045 vs 0.2833 (a 27.8% relative improvement) and WER 0.3256 vs 0.6158 — EasyOCR’s WER is nearly double. The WER gap is larger than the CER gap, which points to EasyOCR compounding character-level slips into whole-word failures on this corpus. Both engines run error-free (error_rate 0.0 on every SROIE and CORD row in the CSV).
Source: summary_metrics.csv — cer and wer columns, sroie_2019 rows. PaddleOCR cer 0.20449 / wer 0.32563; EasyOCR cer 0.28327 / wer 0.61578. Lower is better. 361 samples per engine; both error_rate 0.0.
| Metric (SROIE 2019, n=361) | PaddleOCR | EasyOCR | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.2045 | 0.2833 | summary_metrics.csv · cer, paddleocr/sroie_2019 and easyocr/sroie_2019 rows |
| Word Error Rate (WER) | 0.3256 | 0.6158 | summary_metrics.csv · wer, same rows |
| Error rate (failed pages) | 0.0 | 0.0 | summary_metrics.csv · error_rate, same rows |
Table: summary_metrics.csv — cer / wer / error_rate columns, sroie_2019 rows. Exact values: PaddleOCR cer 0.20449 / wer 0.32563; EasyOCR cer 0.28327 / wer 0.61578. Lower CER/WER is better. Neither engine is the benchmark’s overall accuracy champion — that title belongs to Surya2 (CER 0.1915) and docTR (CER 0.1971) in the same 8-engine run (summary_metrics.csv, surya2 and doctr sroie_2019 rows).
Field Extraction: KIE Through Regex and Through an LLM
Text accuracy ranks engines; field extraction is what downstream systems actually consume. The benchmark’s SROIE field metrics target four flat receipt fields (company, date, address, total) using two postprocessors on each engine’s OCR text: fixed regex patterns (the traditional OCR + rule-based key-information-extraction approach) and an LLM postprocessor (deepseek-v4-flash at temperature 0) with a structured prompt. Through regex, PaddleOCR extracts fields at 0.3254 field F1 to EasyOCR’s 0.1477 — a 2.2× advantage; through the LLM, the gap persists at 0.5810 vs 0.3717 (1.56×).
Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth — 1.0 means every receipt field perfectly recovered, 0 means nothing. The SROIE regex field columns are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text (postprocessed, not native structured output). Note the ranking context: PaddleOCR’s 0.3254 is the third-best regex-field result in the 8-engine benchmark (behind Unlimited-OCR 0.3376 and PaddleOCR-VL 0.3368) and the best among the four purely traditional engines; EasyOCR’s 0.1477 ranks seventh of eight (summary_metrics.csv, field_f1_regex, sroie_2019 rows).
Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| Field extraction (SROIE 2019, n=361) | PaddleOCR | EasyOCR | Source |
|---|---|---|---|
| Field-value F1 (regex) | 0.3254 | 0.1477 | field_method_comparison.csv · regex_field_value_f1, paddleocr/sroie_2019 and easyocr/sroie_2019 rows |
| Field-value F1 (LLM) | 0.5810 | 0.3717 | field_method_comparison.csv · llm_field_value_f1, same rows |
| Field-value accuracy (LLM) | 0.5810 | 0.3712 | field_method_comparison.csv · llm_field_value_accuracy, same rows |
| Docs with all fields exact (LLM) | 0.0748 | 0.0028 | field_method_comparison.csv · llm_document_fields_exact, same rows |
| Median LLM postprocess latency (ms) | 1,817.5 | 2,004.5 | field_method_comparison.csv · llm_median_latency_ms, same rows |
Table: field_method_comparison.csv — regex and llm columns, sroie_2019 rows. The regex columns are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text. LLM postprocessor: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms). “Docs with all fields exact” is the fraction of documents where every target field matched exactly — a much harsher bar than per-field F1; EasyOCR gets all four fields exactly right on 0.28% of receipts.
The EasyOCR LLM Paradox: Mid-Tier Text, Worst Downstream Extraction
This is the page’s most distinctive data point, and it is genuinely counter-intuitive: EasyOCR’s OCR text is mid-pack by character accuracy (SROIE CER 0.2833, fourth of eight engines) — yet when that text is fed to the same LLM postprocessor used for every other engine (deepseek-v4-flash, same prompt, same receipts), its SROIE LLM field F1 of 0.3717 is the lowest of all eight engines in the benchmark — below even Tesseract (0.4389), a CPU engine with a worse CER (0.3347). Only the OCR text changed; the LLM, prompt, and receipts were identical.
Observed pattern, mechanism unverified. A plausible hypothesis — and no more than that — is an output-format convention: how EasyOCR lays out, joins, or separates text lines appears to degrade downstream LLM field extraction for reasons unrelated to raw character accuracy. The benchmark did not isolate this mechanism; the result is documented here as reproducible and stable (EasyOCR’s SROIE row was re-verified in a 2026-08-17 torch 2.8 rerun, r1/r2/r3 byte-identical; the published CSV already carries those corrected values), but no causal claim is made. The practical implication is the opposite of the marketing veneer: on this corpus, choosing EasyOCR for “good enough” text means budgeting for the worst LLM-downstream field recovery of any engine tested.
Source: field_method_comparison.csv — llm_field_value_f1, all eight sroie_2019 rows, 361 samples each (llm_ok_count). LLM postprocessor identical for all engines: deepseek-v4-flash at temperature 0. CER context from summary_metrics.csv, cer column, sroie_2019 rows.
| All 8 engines, SROIE 2019 (n=361 each) | SROIE CER | SROIE LLM field F1 | Source |
|---|---|---|---|
| doctr | 0.1971 | 0.6171 | field_method_comparison.csv · doctr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| surya2 | 0.1915 | 0.6139 | field_method_comparison.csv · surya2/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| unlimited_ocr | 0.6552 | 0.6054 | field_method_comparison.csv · unlimited_ocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| paddleocr_vl_vllm | 0.3370 | 0.5921 | field_method_comparison.csv · paddleocr_vl_vllm/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| PaddleOCR 3.7.0 | 0.2045 | 0.5810 | field_method_comparison.csv · paddleocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| docling | 0.5909 | 0.5685 | field_method_comparison.csv · docling/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| tesseract (CPU) | 0.3347 | 0.4389 | field_method_comparison.csv · tesseract/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| EasyOCR 1.7.2 | 0.2833 | 0.3717 | field_method_comparison.csv · easyocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
Table: field_method_comparison.csv — llm_field_value_f1, all sroie_2019 rows; CER column from summary_metrics.csv, cer, sroie_2019 rows. EasyOCR’s CER (0.2833) ranks fourth of eight — mid-tier text with the worst LLM-downstream field recovery (0.3717, below Tesseract’s 0.4389). The paradox is documented as observed and reproducible; its mechanism is not isolated by this benchmark.
The Operating Envelope: Where EasyOCR Is Genuinely Competitive
Accuracy is not the only axis, and on the operational axis EasyOCR has real, measured counter-advantages. On the same RTX 4090 at the same $0.76/hr, EasyOCR processes SROIE at $0.110 per 1,000 pages vs PaddleOCR’s $0.221 (a 2.0× gap), sustains 124.5 pages/min vs 79.7 (1.56×), and keeps its worst-case tail tight: p95 960.4 ms vs PaddleOCR’s 3,331.4 ms first-page spike (a 3.5× tighter tail). Its footprint is also simpler to deploy — a single PyTorch runtime with wide language coverage, versus PaddlePaddle’s heavier framework stack.
One apparent contradiction deserves an honest explanation: PaddleOCR has the lower median per-page latency (297.0 ms p50 vs EasyOCR’s 413.6 ms) yet the lower pages-per-minute figure (79.7 vs 124.5). The two numbers measure different clocks. Latency p50 is steady-state per-page inference, measured warm (model already loaded); pages/min is wall-clock throughput of the whole run, which includes model initialization and batch effects. EasyOCR’s smaller, faster-loading runtime wins the wall-clock volume race; PaddleOCR’s per-page inference is individually faster once warm. Both numbers are real; they describe different things, and a workload that is dominated by model-load overhead (many small batches, frequent cold starts) will experience EasyOCR’s wall-clock advantage, while a warm long-running pipeline will experience PaddleOCR’s per-page edge.
Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization — the price you would actually pay for the GPU time. Throughput is wall-clock pages per minute including the same initialization. Latency p50/p95 are steady-state per-page inference times measured warm-then-scored (model loading excluded); PaddleOCR’s p95 of 3,331 ms is a first-page/prefill spike, not its steady-state behavior.
Source: summary_metrics.csv — latency_p50_ms / latency_p95_ms columns, sroie_2019 rows. PaddleOCR p50 296.99 / p95 3331.35; EasyOCR p50 413.64 / p95 960.37. Steady-state latency (warm_then_scored measurement mode, excludes model loading).
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. PaddleOCR 0.2214, EasyOCR 0.1098. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026). Neither engine is the benchmark’s cheapest — docTR holds that at $0.048 per 1,000 pages (summary_metrics.csv, doctr/sroie_2019 row).
| Operating envelope (SROIE 2019, n=361) | PaddleOCR | EasyOCR | Source |
|---|---|---|---|
| Latency p50 (ms) | 297.0 | 413.6 | summary_metrics.csv · latency_p50_ms, paddleocr/sroie_2019 and easyocr/sroie_2019 rows |
| Latency p95 (ms) | 3,331.4 | 960.4 | summary_metrics.csv · latency_p95_ms, same rows |
| Pages per minute (wall-clock) | 79.7 | 124.5 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $0.221 | $0.110 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. Both engines GPU (RTX 4090, $0.76/hr price timestamped in manifests); cost includes model init. Exact values: PaddleOCR p50 296.99 / p95 3331.35 / 79.71 pg/min / $0.2214; EasyOCR p50 413.64 / p95 960.37 / 124.53 pg/min / $0.1098. The p50-vs-pages/min inversion is explained in the prose above: steady-state per-page inference (PaddleOCR wins) vs wall-clock throughput including init and batch effects (EasyOCR wins).
CORD (Indonesian Receipts): Both Collapse, But PaddleOCR’s LLM Field Recovery Survives
Neither engine was trained predominantly on Indonesian receipts, so CORD v2 (100 samples, nested fields menu/sub_total/total) functions as a cross-language stress test — and both collapse on raw CER: 0.9083 (PaddleOCR) and 0.9185 (EasyOCR), a language-mismatch wash with both engines effectively unable to read the text. Per the benchmark protocol, CORD numbers are kept quarantined from the SROIE comparison — not merged into any ranking — because CORD’s ground-truth text embeds annotation structure, which inflates raw CER for every engine on top of the genuine language mismatch.
The field metrics tell a different story from the near-identical CERs. Through the LLM postprocessor, PaddleOCR’s field F1 holds at 0.5527 vs EasyOCR’s 0.3378 on CORD — the same SROIE pattern, extended: PaddleOCR’s LLM-downstream recovery withstands the language shock that EasyOCR’s does not, even when both text recognizers fail at the character level. EasyOCR’s CORD LLM field F1 is the second-lowest of the eight engines (above only Tesseract’s 0.1627, summary_metrics.csv/field_method_comparison.csv cord_v2 rows) — again trailing engines with worse CER, the paradox persisting across both datasets. CORD is quoted here for language-robustness context; it is deliberately never pooled with the SROIE numbers into a single leaderboard.
| CORD v2, Indonesian receipts (n=100) | PaddleOCR | EasyOCR | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.9083 | 0.9185 | summary_metrics.csv · cer, paddleocr/cord_v2 and easyocr/cord_v2 rows |
| Field-value F1 (regex) | 0.0154 | 0.0067 | field_method_comparison.csv · regex_field_value_f1, same rows |
| Field-value F1 (LLM) | 0.5527 | 0.3378 | field_method_comparison.csv · llm_field_value_f1, same rows |
| Cost per 1,000 pages | $0.342 | $0.086 | summary_metrics.csv · cost_per_1000_pages, same rows |
| Pages per minute (wall-clock) | 141.0 | 211.8 | summary_metrics.csv · pages_per_minute, same rows |
Table: summary_metrics.csv (cer / cost_per_1000_pages / pages_per_minute) and field_method_comparison.csv (field F1), cord_v2 rows. Do not merge CORD numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth. The regex patterns were written for English formats, which is why regex field F1 collapses to ~0–2% on both engines. PaddleOCR’s LLM field F1 of 0.5527 on CORD repeats its SROIE advantage (0.5810 vs 0.3717) — its LLM-downstream recovery survives the language shock that EasyOCR’s does not.
Who Wins When: The Recap Grid
“Better” is workload-dependent, and this head-to-head splits the axes cleanly: every accuracy axis on English receipts favors PaddleOCR; cost, wall-clock throughput, tail latency, and deployment simplicity favor EasyOCR; and the raw text-accuracy championship belongs to neither (Surya2/docTR) — while EasyOCR’s LLM-downstream result is its biggest caveat, not its selling point.
Frequently Asked Questions
Is PaddleOCR more accurate than EasyOCR on receipts?
Yes, on every accuracy axis measured in this benchmark. On SROIE 2019: CER 0.2045 vs 0.2833 (27.8% relative improvement), WER 0.3256 vs 0.6158, regex field F1 0.3254 vs 0.1477 (2.2×), LLM postprocessed field F1 0.5810 vs 0.3717 (summary_metrics.csv and field_method_comparison.csv, sroie_2019 rows). Neither engine is the benchmark’s overall text champion — Surya2 (CER 0.1915) and docTR (0.1971) hold that title.
Is EasyOCR cheaper than PaddleOCR?
Yes — about 2× cheaper per 1,000 pages on English receipts: $0.110 vs $0.221 on SROIE 2019, widening to $0.086 vs $0.342 on CORD, on the same RTX 4090 at $0.76/hr with cost including model initialization (summary_metrics.csv, cost_per_1000_pages, sroie_2019 and cord_v2 rows). The benchmark’s cheapest engine overall is docTR at $0.048 per 1,000 pages.
Which is faster: PaddleOCR or EasyOCR?
It depends on which clock you mean. PaddleOCR has lower steady-state median latency (297.0 vs 413.6 ms p50), but EasyOCR has higher wall-clock throughput (124.5 vs 79.7 pages/min) because wall-clock pages/min includes model initialization and batch effects and EasyOCR’s lighter runtime loads faster (summary_metrics.csv, latency_p50_ms / pages_per_minute, sroie_2019 rows). For a warm long-running pipeline, PaddleOCR is faster per page; for many small or frequent cold batches, EasyOCR wins the wall-clock race.
Why does EasyOCR have the worst LLM field extraction despite okay character accuracy?
This is the benchmark’s documented paradox, currently without a proven mechanism. EasyOCR’s SROIE CER (0.2833) ranks fourth of eight engines, yet its LLM postprocessed field F1 (0.3717) ranks last — below even Tesseract (0.4389), which has a worse CER. The leading hypothesis is an output-format convention in how EasyOCR lays out or joins text lines that degrades downstream LLM extraction; it is labeled as an observed, reproducible pattern with the mechanism unverified (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows).
Why do both engines score so badly on CORD receipts?
Two compounding causes that the protocol keeps apart from the SROIE ranking: a genuine language mismatch (Indonesian receipts outside both engines’ training focus) and annotation-structure inflation inside CORD’s ground-truth text — CER lands at 0.9083 (PaddleOCR) and 0.9185 (EasyOCR) (summary_metrics.csv, cer, cord_v2 rows). What still separates them is LLM-downstream recovery: PaddleOCR 0.5527 vs EasyOCR 0.3378 field F1 — the SROIE pattern persists even when both recognizers fail at the character level.
Which engine should a receipt pipeline pick, PaddleOCR or EasyOCR?
If your pipeline consumes fields — extracted values for company, date, totals — PaddleOCR is the clear default on receipts: 2.2× field F1 under regex, 1.56× under an LLM, and LLM-downstream recovery that survives the CORD language shock (field_method_comparison.csv). If you need cheap bulk volume, a tight worst-case tail, a lightweight stack, or broad language coverage to get started, EasyOCR is genuinely competitive on the operational axis (2× cheaper, 1.56× wall-clock throughput, 3.5× tighter p95, 80+ languages) — but budget for its weak LLM downstream before committing. These results hold for English and Indonesian receipts on one GPU tier in August 2026; re-run on your target corpus before production decisions (see Limitations).
Where do the numbers on this page come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, regex field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports a head-to-head slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Both engines saw the same images, same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). Both runs completed with error_rate 0.0 on both datasets (summary_metrics.csv error_rate column). The underlying run contains eight engines total; this page compares only the two named engines, with other engines quoted solely as ranking context. The complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.
Runtime Environment
- Hardware: both engines ran on the same NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026).
- Engines: out-of-the-box, no fine-tuning. Versions locked: PaddleOCR 3.7.0 (modern two-stage deep-learning OCR — PP-OCR detection + recognition, GPU) and EasyOCR 1.7.2 (classic ResNet+CRNN with CTC decoding, GPU) — per the public repo model table (README.md) and run manifests. EasyOCR’s SROIE row was re-verified in a 2026-08-17 torch 2.8 rerun (repeat runs r1/r2/r3 byte-identical); the published CSVs carry those corrected values.
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on both engines.
- Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
- Field postprocessing: SROIE regex field metrics are
postprocessed_sroie_receipt_regex_*(field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set. They measure OCR + downstream extraction, not native structured output by either model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended.
Metric Definitions
- CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better.
- WER (Word Error Rate): the same edit-distance calculation at word granularity.
- Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
- Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
- Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
- Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init. They measure different clocks; the p50-vs-pages/min inversion on this page is a measurement-model difference, not an error.
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate, including model initialization.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the paddleocr and easyocr rows here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the paddleocr and easyocr rows here (and to all eight sroie_2019 rows in the paradox table).
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
- Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- Document scope — receipts only: SROIE + CORD. Nothing here measures PaddleOCR’s PP-Structure layout/table/document capabilities, EasyOCR’s 80+ language breadth on non-receipt text, or any other document type. Do not use this page to conclude either engine “wins on everything.”
- Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths should be treated as noise, not engineering truth — though the gaps documented here (27.8% CER, 2.2× regex F1) are far beyond that band.
- Single GPU tier and single price: all numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
- Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the EasyOCR paradox’s magnitude may move with the LLM, though the observed pattern held for this single postprocessor across both datasets. LLM latency (~1,817–2,005 ms median on SROIE, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of either engine’s own latency.
- EasyOCR paradox mechanism unverified: the benchmark documents that EasyOCR’s mid-tier CER text produces the worst LLM-downstream field recovery (0.3717 SROIE / 0.3378 CORD) — an observed, reproducible result under the hypothesis of output text layout conventions, with the causal mechanism explicitly not isolated. Treat it as a measured outcome to plan around, not a proven property of the library.
- Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes.
- CORD CER is not a per-model quality reading: CORD ground truth embeds annotation structure and neither engine was trained predominantly on Indonesian; CORD CER (~0.91) reflects language mismatch + ground-truth inflation. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
- Version pinning: results hold for PaddleOCR 3.7.0 and EasyOCR 1.7.2 (August 2026). Newer releases of either engine may shift every number on this page.
Related references: docTR vs Surya2 Receipt Benchmark · the OCR vs VLM head-to-head · regex rules against LLM extraction · character accuracy vs field accuracy · Receipt OCR Accuracy
Related reading: how AI OCR compares with classic OCR · AI image extraction compared with traditional OCR · AI Document Extraction Pricing (2026)