EasyOCR vs docTR on Receipts
Speed Champion vs Field Extractor (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party head-to-head benchmark · 2 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. Cloud/API OCR services, fine-tuned engines, and the benchmark’s other six engines (Tesseract, PaddleOCR, Docling, Surya2, Unlimited-OCR, PaddleOCR-VL) are out of scope except where quoted as ranking context. The complete 8-engine roundup lives on pixel-level OCR compared with VLM parsing.
Range statement: every number on this page applies to receipts only — SROIE 2019 English receipts and CORD v2 Indonesian receipts. One hardware tier (RTX 4090 at $0.76/hr, price timestamped August 2026), one LLM postprocessor (deepseek-v4-flash at temperature 0), fixed model versions (EasyOCR 1.7.2, docTR v1.0.1). Do not extrapolate these results to other document types, GPUs, or LLMs — the benchmark measures receipt OCR and receipt-field extraction only. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.
On clean English receipts, the modern architecture wins decisively on raw text quality: docTR’s SROIE CER 0.1971 vs EasyOCR’s 0.2833 (30% lower), WER 0.3199 vs 0.6158 (48% lower). Yet run both engines’ text through the same fixed regex patterns and the ranking flips: EasyOCR extracts fields at 1.93× the rate (0.1477 vs 0.0766 field F1) — the benchmark’s "text accuracy ≠ field accuracy" inversion, now between two traditional engines. Add an LLM postprocessor and the flip reverses decisively again: docTR 0.6171 (best of all 8 engines) vs EasyOCR 0.3717 (worst of all 8) — a 1.66× gap and one of the benchmark’s largest LLM-field-F1 deltas. The operating envelope is all docTR: 3.8× faster p50 (108.7 vs 413.6 ms), 3.6× higher throughput (449.3 vs 124.5 pages/min), 2.3× cheaper ($0.048 vs $0.110 per 1,000 pages) — the benchmark’s fastest and cheapest engine on both axes at once.
The trade, in one pair of numbers: docTR reads a receipt page at 108.7 ms p50 for $0.048 per 1,000 pages and, through an LLM postprocessor, extracts fields at 0.6171 F1; EasyOCR reads it at 413.6 ms p50 for $0.110 per 1,000 pages and its LLM-downstream field F1 collapses to 0.3717 — the worst of the eight engines tested. Same receipts, same test split, same RTX 4090. Neither engine "wins"; EasyOCR keeps the regex-field edge and the deployment story, while docTR wins every accuracy, speed, and cost axis measured here.
The two engines represent two generations of deep-learning OCR, both falling on the traditional side of the OCR-versus-VLM divide. EasyOCR (PyTorch-based, 1.7.2) is a classic single-pass CNN + RNN + CTC recognizer — a ResNet feature extractor feeding a sequence model decoded with Connectionist Temporal Classification, with attention-based refinement. Its design center is very wide language and script coverage (80+ languages out of the box) and a famously simple install. docTR (v1.0.1) is a modern two-stage neural pipeline: a detection stage locates text regions, then a recognition stage transcribes them — built around DETR-style transformer detection and a transformer recognizer, engineered for accuracy on printed documents. Character Error Rate (CER) measures insertions, deletions, and substitutions divided by ground-truth characters — a CER of 0.197 means ~19.7 characters misread per 100; Word Error Rate (WER) applies the same edit-distance logic at whole-word granularity. Lower is better on both. Neither engine is a vision-language model (VLM) — both output raw text, not understood structure.
Text Accuracy on SROIE (English Receipts): docTR’s Clear Edge
On the 361 English receipts of the SROIE 2019 test split, the modern architecture wins on both text metrics: CER 0.1971 vs 0.2833 (30% relative improvement) and WER 0.3199 vs 0.6158 — EasyOCR’s WER is nearly double. The WER gap (48%) is much larger than the CER gap (30%), which points to EasyOCR compounding character-level slips into whole-word failures on this corpus. Both engines run error-free (error_rate 0.0 on every SROIE and CORD row in the CSV). Neither engine is the benchmark’s overall CER champion — that title belongs to Surya2 (0.1915) and docTR itself sits second; EasyOCR ranks fourth of eight.
Source: summary_metrics.csv — cer and wer columns, sroie_2019 rows. docTR cer 0.19707 / wer 0.31990; EasyOCR cer 0.28327 / wer 0.61578. Lower is better. 361 samples per engine; both error_rate 0.0.
| Metric (SROIE 2019, n=361) | docTR | EasyOCR | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.1971 | 0.2833 | summary_metrics.csv · cer, doctr/sroie_2019 and easyocr/sroie_2019 rows |
| Word Error Rate (WER) | 0.3199 | 0.6158 | summary_metrics.csv · wer, same rows |
| Error rate (failed pages) | 0.0 | 0.0 | summary_metrics.csv · error_rate, same rows |
Table: summary_metrics.csv — cer / wer / error_rate columns, sroie_2019 rows. Exact values: docTR cer 0.19707 / wer 0.31990; EasyOCR cer 0.28327 / wer 0.61578. Relative gaps: CER 30% lower, WER 48% lower for docTR. Lower CER/WER is better. CER ranking context within the same 8-engine run: Surya2 0.1915, docTR 0.1971, PaddleOCR 0.2045, EasyOCR 0.2833 (summary_metrics.csv, cer, sroie_2019 rows).
The Regex-Field Inversion: Worse Text, More Recoverable Fields
Benchmark both engines’ raw text through the same fixed regex patterns on the four SROIE receipt fields (company, date, address, total) — the traditional OCR + rule-based key-information-extraction (KIE) approach — and the ranking flips: EasyOCR extracts fields at 0.1477 field F1 to docTR’s 0.0766, a 1.93× advantage for the engine with the worse character accuracy. This is the benchmark’s recurring "text accuracy ≠ field accuracy" inversion — the same pattern seen between traditional OCR and document-parsing VLMs on docTR vs Surya2 — now occurring between two traditional engines that output the same kind of raw line text.
Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth: 1.0 means every receipt field perfectly recovered, 0 means nothing. The mechanism behind the flip is a property of the regex pattern set, not of recognition quality per se: the patterns were written once per dataset for formatted values like RM 12.00 or 14/08/2020. docTR’s clean-but-raw line text — accurate by CER, but preserving original casing and separator noise — defeats the fixed patterns; EasyOCR’s output happens to match them more often. The “regex field extraction” columns are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: they measure OCR text + downstream rule-based extraction, not native structured output. A per-format, heavily tuned pattern library could score differently for either engine — the pattern set is a fixed measurement instrument, not a tuned production parser.
Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| Regex postprocessing (SROIE 2019, n=361) | docTR | EasyOCR | Source |
|---|---|---|---|
| Field-value F1 (regex) | 0.0766 | 0.1477 | field_method_comparison.csv · regex_field_value_f1, doctr/sroie_2019 and easyocr/sroie_2019 rows |
| Field-value accuracy (regex) | 0.0623 | 0.1267 | field_method_comparison.csv · regex_field_value_accuracy, same rows |
| Document-fields exact (regex) | 0.0000 | 0.0000 | field_method_comparison.csv · regex_document_fields_exact, same rows |
Table: field_method_comparison.csv — regex columns, sroie_2019 rows. These are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text (postprocessed, not native extraction). docTR’s regex field F1 of 0.0766 is the second-lowest of all eight engines in the underlying run despite the second-best CER — the pattern set was written once per dataset, and docTR’s clean-but-raw line text is not regex-friendly on these four fields.
The LLM Lever: The Flip Reverses, Decisively
Feed both engines’ OCR text to an LLM postprocessor (deepseek-v4-flash at temperature 0) with a structured extraction prompt, and the field ranking reverses back — with the widest margin of any pairing in the benchmark: docTR 0.6171 vs EasyOCR 0.3717 field F1, a 1.66× gap. docTR’s result is the highest LLM field F1 of all eight engines; EasyOCR’s is the lowest. Where the regex pattern set punished docTR’s clean text, the LLM rewards it — and EasyOCR’s mid-tier text, which happened to be regex-friendly, degrades under the same prompt.
This is the same LLM-convergence pattern seen across the full eight-engine benchmark — LLM postprocessing pulls healthy engines into a 0.57–0.62 field-F1 band because it understands semantics (numbers, dates, names) instead of matching character shapes — with EasyOCR as the stark exception. The lever is not free: an LLM call adds roughly 2.0–2.1 s of median latency per document on top of OCR time (1,996.3 ms for docTR’s text, 2,004.5 ms for EasyOCR’s, API-incurred and identical in kind), and it does not rescue text an engine fundamentally failed to read. But for this pairing the postprocessor becomes the decisive component: with an LLM in the pipeline, the docTR choice compounds.
| LLM postprocessing (SROIE 2019, n=361) | docTR | EasyOCR | Source |
|---|---|---|---|
| Field-value F1 (LLM) | 0.6171 | 0.3717 | field_method_comparison.csv · llm_field_value_f1, doctr/sroie_2019 and easyocr/sroie_2019 rows |
| Field-value accuracy (LLM) | 0.6170 | 0.3712 | field_method_comparison.csv · llm_field_value_accuracy, same rows |
| Document-fields exact (LLM) | 0.1496 | 0.0028 | field_method_comparison.csv · llm_document_fields_exact, same rows |
| Median LLM postprocess latency (ms) | 1,996.3 | 2,004.5 | field_method_comparison.csv · llm_median_latency_ms, same rows |
Table: field_method_comparison.csv — llm_* columns, sroie_2019 rows. LLM model: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms). “Document-fields exact” is the fraction of documents where every target field matched exactly — a much harsher bar than per-field F1; EasyOCR gets all four fields exactly right on 0.28% of receipts.
The EasyOCR LLM Paradox: Mid-Tier Text, Worst Downstream Extraction
This head-to-head’s most counter-intuitive data point — first documented on PaddleOCR vs EasyOCR and confirmed here against a different opponent: EasyOCR’s OCR text is mid-pack by character accuracy (SROIE CER 0.2833, fourth of eight engines) — yet when that text is fed to the same LLM postprocessor used for every other engine (deepseek-v4-flash, same prompt, same receipts), its SROIE LLM field F1 of 0.3717 is the lowest of all eight engines in the benchmark — below even Tesseract (0.4389), a CPU engine with a worse CER (0.3347). Only the OCR text changed; the LLM, prompt, and receipts were identical.
Observed pattern, mechanism unverified. A plausible hypothesis — and no more than that — is an output-format convention: how EasyOCR lays out, joins, or separates text lines appears to degrade downstream LLM field extraction for reasons unrelated to raw character accuracy. The benchmark did not isolate this mechanism; the result is documented here as reproducible and stable (EasyOCR’s SROIE row was re-verified in a 2026-08-17 torch 2.8 rerun, r1/r2/r3 byte-identical; the published CSV already carries those corrected values), but no causal claim is made. docTR’s cleaner base text + the LLM lands at the top of the same band EasyOCR misses.
Source: field_method_comparison.csv — llm_field_value_f1, all eight sroie_2019 rows, 361 samples each (llm_ok_count). LLM postprocessor identical for all engines: deepseek-v4-flash at temperature 0. CER context from summary_metrics.csv, cer column, sroie_2019 rows.
| All 8 engines, SROIE 2019 (n=361 each) | SROIE CER | SROIE LLM field F1 | Source |
|---|---|---|---|
| docTR v1.0.1 | 0.1971 | 0.6171 | field_method_comparison.csv · doctr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| surya2 | 0.1915 | 0.6139 | field_method_comparison.csv · surya2/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| unlimited_ocr | 0.6552 | 0.6054 | field_method_comparison.csv · unlimited_ocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| paddleocr_vl_vllm | 0.3370 | 0.5921 | field_method_comparison.csv · paddleocr_vl_vllm/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| paddleocr | 0.2045 | 0.5810 | field_method_comparison.csv · paddleocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| docling | 0.5909 | 0.5685 | field_method_comparison.csv · docling/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| tesseract (CPU) | 0.3347 | 0.4389 | field_method_comparison.csv · tesseract/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
| EasyOCR 1.7.2 | 0.2833 | 0.3717 | field_method_comparison.csv · easyocr/sroie_2019 row, llm_field_value_f1 (cer: summary_metrics.csv) |
Table: field_method_comparison.csv — llm_field_value_f1, all sroie_2019 rows; CER column from summary_metrics.csv, cer, sroie_2019 rows. docTR’s 0.6171 is the highest LLM field F1 in the benchmark; EasyOCR’s CER (0.2833) ranks fourth of eight — mid-tier text with the worst LLM-downstream field recovery (0.3717, below Tesseract’s 0.4389). The paradox is documented as observed and reproducible; its mechanism is not isolated by this benchmark.
The Operating Envelope: docTR Is the Benchmark’s Speed AND Cost Champion
Accuracy decides which engine reads best; the operating envelope decides which one finishes. On the same RTX 4090 at the same recorded $0.76/hr rate, docTR sustains 449.3 pages/min at 108.7 ms p50 per page for $0.048 per 1,000 pages; EasyOCR sustains 124.5 pages/min at 413.6 ms p50 for $0.110 per 1,000 pages — a 3.8× latency gap, a 3.6× throughput gap, and a 2.3× cost gap, all in docTR’s favor. Across the full eight-engine run, docTR’s 108.7 ms p50, 449.3 pages/min, and $0.048 cost are each the best of any engine measured — docTR is simultaneously the benchmark’s fastest and cheapest engine.
Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization — the price you would actually pay for the GPU time. Throughput is wall-clock pages per minute including that same initialization. Latency p50/p95 are steady-state per-page inference times measured warm-then-scored (model loading excluded); EasyOCR’s tail is proportionally worse — 960.4 ms p95 against docTR’s 281.4 ms, a 3.4× gap. EasyOCR is still genuinely cheaper than most of the other engines measured (its $0.110 is the second-lowest per-1,000-pages figure in the benchmark, behind only docTR) — it is mid-cost and mid-speed, not expensive or slow.
Source: summary_metrics.csv — latency_p50_ms / latency_p95_ms columns, sroie_2019 rows. docTR p50 108.72 / p95 281.38; EasyOCR p50 413.64 / p95 960.37. Steady-state latency (warm_then_scored measurement mode, excludes model loading).
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. docTR 0.0479, EasyOCR 0.1098. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026). docTR’s $0.048 is the lowest per-1,000-pages cost of any engine in the benchmark; EasyOCR’s $0.110 is the second-lowest (summary_metrics.csv, cost_per_1000_pages, all sroie_2019 rows).
| Operating envelope (SROIE 2019, n=361) | docTR | EasyOCR | Source |
|---|---|---|---|
| Latency p50 (ms) | 108.7 | 413.6 | summary_metrics.csv · latency_p50_ms, doctr/sroie_2019 and easyocr/sroie_2019 rows |
| Latency p95 (ms) | 281.4 | 960.4 | summary_metrics.csv · latency_p95_ms, same rows |
| Pages per minute (wall-clock) | 449.3 | 124.5 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $0.048 | $0.110 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. Both engines GPU (RTX 4090, $0.76/hr price timestamped in manifests); cost includes model init, not pure steady-state throughput. Exact values: docTR p50 108.72 / p95 281.38 / 449.31 pg/min / $0.0479; EasyOCR p50 413.64 / p95 960.37 / 124.53 pg/min / $0.1098. Benchmark-wide bests: docTR holds the lowest p50 latency, highest pages/min, and lowest cost of all eight engines (summary_metrics.csv, sroie_2019 rows).
CORD (Indonesian Receipts): Both Collapse on Text, the LLM Gap Widens
Neither engine was trained predominantly on Indonesian receipts, so CORD v2 (100 samples, nested fields menu/sub_total/total) functions as a cross-language stress test — and both collapse on raw CER: 0.9101 (docTR) and 0.9185 (EasyOCR), a language-mismatch wash with both engines effectively unable to read the text. Per the benchmark protocol, CORD numbers are kept quarantined from the SROIE comparison — not merged into any ranking — because CORD’s ground-truth text embeds annotation structure, which inflates raw CER for every engine on top of the genuine language mismatch.
The field metrics show the SROIE pattern extending — and widening. Through the LLM postprocessor, docTR’s field F1 holds at 0.5500 vs EasyOCR’s 0.3378 on CORD — a 1.63× gap, the same ordering as SROIE’s 1.66×, even when both text recognizers fail at the character level. Through regex patterns, docTR recovers no fields (0.0000 field F1 — a literal zero in the CSV, not a missing value) because the English-format patterns matched nothing in Indonesian text, while EasyOCR scrapes 0.0067. docTR’s cost edge also narrows and reverses on CORD ($0.094 vs EasyOCR’s $0.086 per 1,000 pages) — but its wall-clock throughput advantage grows to 500.4 vs 211.8 pages/min (2.4×). CORD is quoted here for language-robustness context; it is deliberately never pooled with the SROIE numbers into a single leaderboard.
| CORD v2, Indonesian receipts (n=100) | docTR | EasyOCR | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.9101 | 0.9185 | summary_metrics.csv · cer, doctr/cord_v2 and easyocr/cord_v2 rows |
| Field-value F1 (regex) | 0.0000 | 0.0067 | field_method_comparison.csv · regex_field_value_f1, same rows |
| Field-value F1 (LLM) | 0.5500 | 0.3378 | field_method_comparison.csv · llm_field_value_f1, same rows |
| Pages per minute (wall-clock) | 500.4 | 211.8 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $0.094 | $0.086 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv (cer / pages_per_minute / cost_per_1000_pages) and field_method_comparison.csv (field F1), cord_v2 rows. Do not merge CORD numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth. The regex patterns were written for English formats, which is why regex field F1 collapses to ~0–1% on both engines; docTR’s 0.0000 is a literal zero recorded in the CSV, not a missing value. The LLM field-F1 gap (0.5500 vs 0.3378) extends the SROIE pattern across languages; the cost order reverses (EasyOCR $0.086 vs docTR $0.094) while the throughput gap widens (2.4×).
Who Wins When: The Recap Grid
“Better” is workload-dependent, and this head-to-head splits the axes cleanly: raw text accuracy, LLM-downstream fields, speed, throughput, and cost all favor docTR; fixed-regex field extraction and the deployment story favor EasyOCR — with the caveat that EasyOCR’s LLM-downstream result is its biggest risk, not its selling point.
Frequently Asked Questions
Is docTR more accurate than EasyOCR on receipts?
Yes, on every raw-text and field-extraction accuracy axis in this benchmark. On SROIE 2019: CER 0.1971 vs 0.2833 (30% lower), WER 0.3199 vs 0.6158 (48% lower), LLM postprocessed field F1 0.6171 vs 0.3717 (1.66×, best of 8 vs worst of 8) — but not on regex field F1, where EasyOCR wins 0.1477 vs 0.0766 (summary_metrics.csv and field_method_comparison.csv, sroie_2019 rows). Neither engine is the benchmark’s overall CER champion — Surya2 (0.1915) holds that title by a hair over docTR.
Why does docTR have better text accuracy but worse regex field extraction than EasyOCR?
Because the two metrics grade different outputs against a fixed measurement instrument. docTR returns clean raw line text — accurate by CER, but preserving original casing and separators — and the fixed regex patterns, written once per dataset for formatted values, mostly fail against it: SROIE regex field F1 0.0766 (field_method_comparison.csv, regex_field_value_f1, doctr/sroie_2019 row). EasyOCR’s output happens to match the patterns at 0.1477. Feed both to an LLM instead and the gap reverses to 1.66× in docTR’s favor — the regex pattern set, not the OCR, was the bottleneck.
Why does EasyOCR have the worst LLM field extraction despite okay character accuracy?
This is the benchmark’s documented paradox, currently without a proven mechanism. EasyOCR’s SROIE CER (0.2833) ranks fourth of eight engines, yet its LLM postprocessed field F1 (0.3717) ranks last — below even Tesseract (0.4389), which has a worse CER. The leading hypothesis is an output-format convention in how EasyOCR lays out or joins text lines that degrades downstream LLM extraction; it is labeled as an observed, reproducible pattern with the mechanism unverified (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows). The same paradox is documented against a different opponent on PaddleOCR vs EasyOCR.
How much faster and cheaper is docTR than EasyOCR?
3.8× lower p50 latency (108.7 vs 413.6 ms), 3.4× lower p95 (281.4 vs 960.4 ms), 3.6× higher throughput (449.3 vs 124.5 pages/min), and 2.3× lower cost per 1,000 pages ($0.048 vs $0.110) on the same RTX 4090 at $0.76/hr (summary_metrics.csv, latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows). docTR is the benchmark’s fastest and cheapest engine on all three clocks; EasyOCR is mid-cost (second-cheapest) and mid-speed (second-fastest throughput).
Why do both engines score so badly on CORD receipts?
Two compounding causes that the benchmark protocol keeps apart from the SROIE ranking: a genuine language mismatch (Indonesian receipts outside both engines’ training focus) and annotation-structure inflation inside CORD’s ground-truth text — CER lands at 0.9101 (docTR) and 0.9185 (EasyOCR) (summary_metrics.csv, cer, cord_v2 rows). What still separates them is LLM-downstream recovery: docTR 0.5500 vs EasyOCR 0.3378 field F1 — the SROIE pattern persists, and widens, even when both recognizers fail at the character level.
Which engine should a receipt pipeline pick, EasyOCR or docTR?
For a pipeline whose goal is extracted fields at volume with metered cost, docTR dominates on this corpus: better raw text (30% lower CER), the best LLM-downstream fields of any engine (0.6171 vs 0.3717), and a 2.3× cost advantage, 3.6× throughput, 3.8× latency — all four axes at once (summary_metrics.csv / field_method_comparison.csv, sroie_2019 rows). EasyOCR remains a legitimate choice for cheap, easy-to-install, multi-script bulk OCR on clean documents where regex extraction or raw text at modest volume is the task and LLM-downstream field quality matters less — but budget for its measured weak LLM downstream before committing. These results hold for English and Indonesian receipts on one GPU tier in August 2026; re-run on your target corpus before production decisions (see Limitations).
Where do the numbers on this page come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, regex field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports a head-to-head slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Both engines saw the same images, same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). Both runs completed with error_rate 0.0 on both datasets (summary_metrics.csv error_rate column). The underlying run contains eight engines total; this page compares only the two named engines, with other engines quoted solely as ranking context. The complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.
Runtime Environment
- Hardware: both engines ran on the same NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026).
- Engines: out-of-the-box, no fine-tuning. Versions locked: EasyOCR 1.7.2 (classic CNN + RNN + CTC recognizer, ResNet feature extractor, GPU) and docTR v1.0.1 (modern two-stage neural OCR — DETR-style transformer detection + recognition, GPU) — per the public repo model table (README.md) and run manifests. EasyOCR’s SROIE row was re-verified in a 2026-08-17 torch 2.8 rerun (repeat runs r1/r2/r3 byte-identical); the published CSVs carry those corrected values.
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on both engines.
- Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
- Field postprocessing: SROIE regex field metrics are
postprocessed_sroie_receipt_regex_*(field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set written once per dataset. They measure OCR + downstream extraction, not native structured output by either model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended.
Metric Definitions
- CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better.
- WER (Word Error Rate): the same edit-distance calculation at word granularity.
- Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
- Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
- Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
- Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init. They measure different clocks.
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate, including model initialization.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the easyocr and doctr rows here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the easyocr and doctr rows here (and to all eight sroie_2019 rows in the paradox table).
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
- Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- Document scope — receipts only: SROIE + CORD. Nothing here measures EasyOCR’s 80+ language breadth on non-receipt text, docTR’s behavior on tables/forms/long documents, or any other document type. Do not use this page to conclude either engine “wins on everything.”
- Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths should be treated as noise, not engineering truth — though the gaps documented here (30% CER, 48% WER, 1.66× LLM F1) are far beyond that band.
- Single GPU tier and single price: all numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
- Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the EasyOCR paradox’s magnitude may move with the LLM, though the observed pattern held for this single postprocessor across both datasets. LLM latency (~2.0–2.1 s median on SROIE, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of either engine’s own latency.
- EasyOCR paradox mechanism unverified: the benchmark documents that EasyOCR’s mid-tier CER text produces the worst LLM-downstream field recovery (0.3717 SROIE / 0.3378 CORD) — an observed, reproducible result under the hypothesis of output text layout conventions, with the causal mechanism explicitly not isolated. Treat it as a measured outcome to plan around, not a proven property of the library.
- Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes. docTR’s regex-field disadvantage (0.0766 vs 0.1477) is a property of this fixed instrument, not a claim about what a tuned parser could recover.
- CORD CER is not a per-model quality reading: CORD ground truth embeds annotation structure and neither engine was trained predominantly on Indonesian; CORD CER (~0.91) reflects language mismatch + ground-truth inflation. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
- Two engines only: this head-to-head deliberately excludes the other six engines of the underlying run, cloud/API OCR services, and hosted VLM APIs; their latency and pricing models differ fundamentally from the local engines measured here.
- Version pinning: results hold for EasyOCR 1.7.2 and docTR v1.0.1 (August 2026). Newer releases of either engine may shift every number on this page.
Related references: PaddleOCR vs EasyOCR Receipt Benchmark · docTR vs Surya2 Receipt Benchmark · Traditional OCR vs Document Parsing VLMs · pattern matching vs model-based field extraction · measuring accuracy per field, not per character · OCR Cost per 1,000 Pages
Related reading: why AI OCR accuracy diverges from traditional OCR · why AI extraction beats OCR on images · AI Document Extraction Pricing (2026)