docTR vs Surya2: Receipt OCR
Head-to-Head Benchmark (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party head-to-head benchmark · 2 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no tables, forms, invoices, contracts, or long documents. Surya2’s marketed strengths (layout analysis, table recognition, long documents, arbitrary languages) are not measured here. Cloud/API OCR services, other open-source engines (only these two are compared), fine-tuned models, and full-text metrics outside CER/WER are out of scope. The complete 8-engine roundup lives on OCR engines vs vision-language models.
Scope of every number on this page: receipts (SROIE 2019 English, CORD v2 Indonesian), one GPU tier (RTX 4090 at $0.76/hr), August 2026 model versions. Do not extrapolate these results to invoices, tables, or complex layouts — the benchmark measures receipt OCR and receipt-field extraction only. All figures come from the benchmark’s results/summary_metrics.csv and results/field_method_comparison.csv, mirrored in the public GitHub repository and cited row-by-row.
The two best text recognizers in the benchmark are statistically tied on receipt character accuracy — one traditional OCR engine and one document-parsing VLM: SROIE 2019 CER 0.1971 (docTR) vs 0.1915 (Surya2), a 0.006-point gap. They diverge only when you ask them to produce fields: through fixed regex patterns, Surya2’s case-folded, label-structured output extracts fields at 4.2× the rate of docTR’s clean-but-raw line text (0.3183 vs 0.0766 field F1). Add an LLM postprocessor and the gap nearly disappears — docTR 0.6171 vs Surya2 0.6139 field F1. The real, deciding difference between these two engines is the operating envelope: 24.5× latency, 37× throughput, and 22× cost per 1,000 pages, all in docTR’s favor on identical hardware.
The trade, in one pair of numbers: docTR reads a receipt page at 108.7 ms p50 for $0.048 per 1,000 pages; Surya2 reads it at 2,668.0 ms p50 for $1.061 per 1,000 pages — same receipts, same test split, same RTX 4090. Neither engine "wins"; they win on different axes, and the point of this page is to show both axes from the same controlled run.
Character Error Rate (CER) is the classic OCR yardstick: insertions, deletions, and substitutions divided by ground-truth characters — a CER of 0.197 means roughly 19.7 characters misread per 100. Word Error Rate (WER) applies the same edit-distance calculation at word granularity. This is the dimension where the two engines are inseparable.
On SROIE 2019, docTR and Surya2 are the benchmark’s two strongest character recognizers, separated by 0.006 points — inside the noise band for a 361-sample corpus. WER tells the same story: Surya2 0.2735, docTR 0.3199. The "VLMs beat traditional OCR" narrative does not survive this comparison on raw character accuracy of English receipts.
Character Accuracy: A Statistical Tie
The tie matters because the two engines are architecturally opposite. docTR is a traditional neural OCR pipeline in two stages: a detector localizes words and a recognizer transcribes them, preserving the raw case and layout of the printed text. Surya2 is a document-parsing vision-language model (VLM): it reads the whole document image and outputs understood text — case-folded, label/value pairs merged, lines reordered — which is closer to what downstream systems want but farther from exact character matches. CER scores exact character matches, so it is mildly conservative against Surya2; the fact that the tie survives despite that asymmetry is what makes it meaningful.
| Metric (SROIE 2019, n=361) | docTR | Surya2 | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.1971 | 0.1915 | summary_metrics.csv · cer, doctr/sroie_2019 and surya2/sroie_2019 rows |
| Word Error Rate (WER) | 0.3199 | 0.2735 | summary_metrics.csv · wer, same rows |
Table: summary_metrics.csv — cer and wer columns, sroie_2019 rows. Exact values: docTR cer 0.19707 / wer 0.31990; Surya2 cer 0.19147 / wer 0.27352. Lower is better; both runs completed with error_rate 0.0.
Where They Diverge: Out-of-Box Field Extraction Flips the Outcome
Benchmark the engines’ raw text through the same fixed regex patterns on the four SROIE receipt fields (company, date, address, total) — the traditional OCR + rule-based key-information-extraction (KIE) approach — and the ranking flips: Surya2 extracts fields at 0.3183 field F1 to docTR’s 0.0766, a 4.2× advantage. docTR has the best-class character accuracy and the worst regex field extraction in the eight-engine run — text accuracy and field accuracy are decoupled.
Field-value F1 is the harmonic mean of precision and recall over extracted field values against ground truth: 1.0 means every receipt field perfectly recovered, 0 means nothing. The mechanism behind the flip is the output shape difference described above. Regular expressions were written for formatted values like RM 12.00 or 14/08/2020; Surya2’s normalized, label-structured output (case-folded, key-value merged) matches those patterns far more often, while docTR’s raw line text — accurate by CER, but with original casing and separator noise — defeats them. The "regex field extraction" columns are the benchmark’s postprocessed_sroie_receipt_regex_* metrics: they measure OCR text + downstream rule-based extraction, not native structured output.
Source: field_method_comparison.csv — regex_field_value_f1 / llm_field_value_f1 columns, sroie_2019 rows (0–1 stored decimals shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| Regex postprocessing (SROIE 2019, n=361) | docTR | Surya2 | Source |
|---|---|---|---|
| Field-value F1 (regex) | 0.0766 | 0.3183 | field_method_comparison.csv · regex_field_value_f1, doctr/sroie_2019 and surya2/sroie_2019 rows |
| Field-value accuracy (regex) | 0.0623 | 0.2999 | field_method_comparison.csv · regex_field_value_accuracy, same rows |
| Document-fields exact (regex) | 0.0000 | 0.0194 | field_method_comparison.csv · regex_document_fields_exact, same rows |
Table: field_method_comparison.csv — regex columns, sroie_2019 rows. These are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text (postprocessed, not native extraction). docTR’s regex field F1 of 0.0766 is the lowest of all eight engines in the underlying run despite the second-best CER.
The LLM Lever: Engine Choice Stops Mattering
Feed both engines’ OCR text to an LLM postprocessor (deepseek-v4-flash, temperature 0) with a structured extraction prompt, and the field gap nearly vanishes: docTR 0.6171 vs Surya2 0.6139 field F1 — a 0.003-point edge for the traditional engine, a tie in practice. The postprocessor, not the OCR engine, becomes the deciding component.
This is the same pattern seen across the full eight-engine benchmark: LLM postprocessing pulls healthy engines into a convergent field-F1 band because it understands semantics (numbers, dates, names) instead of matching character shapes. docTR’s cleaner base text edges ahead by a hair; Surya2’s normalized structure loses that tiny advantage under the LLM. Two costs come with the lever: an LLM call adds roughly 2.0–2.3 s of median latency per document on top of OCR time (1,996.3 ms for docTR’s text, 2,261.7 ms for Surya2’s, API-incurred and identical in kind), and it does not rescue text an engine fundamentally failed to read.
| LLM postprocessing (SROIE 2019, n=361) | docTR | Surya2 | Source |
|---|---|---|---|
| Field-value F1 (LLM) | 0.6171 | 0.6139 | field_method_comparison.csv · llm_field_value_f1, doctr/sroie_2019 and surya2/sroie_2019 rows |
| Field-value accuracy (LLM) | 0.6170 | 0.6136 | field_method_comparison.csv · llm_field_value_accuracy, same rows |
| Document-fields exact (LLM) | 0.1496 | 0.1551 | field_method_comparison.csv · llm_document_fields_exact, same rows |
| Median LLM latency (ms) | 1,996.3 | 2,261.7 | field_method_comparison.csv · llm_median_latency_ms, same rows |
Table: field_method_comparison.csv — llm_* columns, sroie_2019 rows. LLM model: deepseek-v4-flash at temperature 0 (llm_model column). LLM latency is API-incurred and separate from engine latency (summary_metrics.csv latency_p50_ms).
The Operating Envelope: Where the Real Difference Lives
Character accuracy ties, field accuracy converges under an LLM — but a batch pipeline does not care about either if the numbers don’t finish. On the same RTX 4090 at the same recorded $0.76/hr rate, docTR sustains 449.3 pages/min at 108.7 ms p50 per page for $0.048 per 1,000 pages; Surya2 sustains 12.1 pages/min at 2,668.0 ms p50 for $1.061 per 1,000 pages — a 37× throughput gap, a 24.5× latency gap, and a 22.2× cost gap. A pipeline sized for Surya2’s pace is a different architecture conversation than one sizing docTR’s pace.
Cost is computed as wall-clock runtime × the RunPod RTX 4090 rate ($0.76/hour, price timestamped in the run manifests), including model initialization — the price you would actually pay for the GPU time. Throughput is wall-clock pages per minute including that same initialization. Latency p50/p95 are steady-state per-page inference times measured warm-then-scored (model loading excluded); Surya2’s tail is proportionally worse — 5,872.2 ms p95 against docTR’s 281.4 ms — because VLM prefill/decode spikes dominate the tail on first pages.
Source: summary_metrics.csv — latency_p50_ms column, sroie_2019 rows. docTR 108.7166, Surya2 2667.9800. Steady-state latency (warm_then_scored measurement mode).
Source: summary_metrics.csv — cost_per_1000_pages column, sroie_2019 rows. docTR 0.0479, Surya2 1.0609. Cost = wall-clock runtime × $0.76/hr including model init, price timestamped in run manifests (August 2026).
| Operating envelope (SROIE 2019, n=361) | docTR | Surya2 | Source |
|---|---|---|---|
| Latency p50 (ms) | 108.7 | 2,668.0 | summary_metrics.csv · latency_p50_ms, doctr/sroie_2019 and surya2/sroie_2019 rows |
| Latency p95 (ms) | 281.4 | 5,872.2 | summary_metrics.csv · latency_p95_ms, same rows |
| Pages per minute | 449.3 | 12.1 | summary_metrics.csv · pages_per_minute, same rows |
| Cost per 1,000 pages | $0.048 | $1.061 | summary_metrics.csv · cost_per_1000_pages, same rows |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows. Both engines GPU (RTX 4090, $0.76/hr price timestamped in manifests); cost includes model init, not pure steady-state throughput. Exact values: docTR p50 108.7 / p95 281.4 / 449.3 pg/min / $0.0479; Surya2 p50 2668.0 / p95 5872.2 / 12.1 pg/min / $1.0609.
CORD (Indonesian Receipts): Both Engines Collapse, Quarantined by Protocol
Neither engine was trained predominantly on Indonesian receipts, so CORD v2 (100 samples, nested fields menu/sub_total/total) functions as a cross-language stress test — and both collapse: CER 0.8959 (Surya2) and 0.9101 (docTR). Per the benchmark protocol, CORD numbers are kept quarantined from the SROIE comparison — they are not merged into any ranking — because CORD’s ground-truth text embeds annotation structure, which inflates raw CER for every engine on top of the genuine language mismatch.
On the field metrics the same divergence pattern from SROIE holds, compressed: through regex patterns docTR recovers no fields (0.0000 field F1 — a literal zero in the CSV, not a missing value) because the English-format patterns matched nothing in Indonesian text, while Surya2’s normalized output scrapes 0.2458. The LLM then re-converges the pair to 0.5500 (docTR) and 0.5203 (Surya2) — the language shock is absorbed by the postprocessor, not by the engine. CORD is quoted here for language-robustness context; it is deliberately never pooled with the SROIE numbers into a single leaderboard.
| CORD v2, Indonesian receipts (n=100) | docTR | Surya2 | Source |
|---|---|---|---|
| Character Error Rate (CER) | 0.9101 | 0.8959 | summary_metrics.csv · cer, doctr/cord_v2 and surya2/cord_v2 rows |
| Field-value F1 (regex) | 0.0000 | 0.2458 | field_method_comparison.csv · regex_field_value_f1, same rows |
| Field-value F1 (LLM) | 0.5500 | 0.5203 | field_method_comparison.csv · llm_field_value_f1, same rows |
Table: summary_metrics.csv (cer) and field_method_comparison.csv (field F1), cord_v2 rows. Do not merge these numbers into any SROIE ranking: CORD CER combines genuine language mismatch with annotation-structure inflation in the ground truth; the regex patterns were written for English formats. docTR’s regex field F1 of 0.0000 is a literal zero recorded in the CSV, not a missing value.
Who Wins When: The Recap Grid
"Better" is workload-dependent. These two engines split the axes cleanly, and the split is the finding: character accuracy ties, structured-field convenience favors Surya2, every cost/latency/throughput axis favors docTR, and an LLM postprocessor makes engine choice nearly irrelevant for final field quality.
Frequently Asked Questions
Is Surya2 more accurate than docTR on receipts?
No — on raw character accuracy they are statistically tied: SROIE CER 0.1915 (Surya2) vs 0.1971 (docTR), a 0.006-point gap (summary_metrics.csv, cer, sroie_2019 rows). Where they differ is structured field output out-of-box (Surya2 wins 4.2× under regex) and the operating envelope (docTR wins 24.5× on latency, 22.2× on cost, 37× on throughput).
Why does docTR have the best character accuracy but the worst regex field extraction?
Because the two metrics grade different outputs. docTR returns clean raw line text — preserving original case and separators — and the fixed regex patterns, written for formatted values, mostly fail against it: its SROIE regex field F1 is 0.0766, the lowest of all eight engines in the underlying run, against the second-best CER (0.1971) (summary_metrics.csv cer, field_method_comparison.csv regex_field_value_f1). Surya2’s case-folded, label-structured output happens to match the patterns at 0.3183. Feed both to an LLM instead and the gap collapses to 0.003 — the regex, not the OCR, was the bottleneck.
Does adding an LLM postprocessor make docTR and Surya2 equal?
Nearly exactly — docTR 0.6171 vs Surya2 0.6139 LLM field F1 on SROIE (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows). The cost of that convergence is an extra ~2.0–2.3 s of median LLM latency per document (llm_median_latency_ms, same rows), suited to asynchronous batch processing rather than synchronous per-page waits.
How much faster and cheaper is docTR than Surya2?
24.5× lower p50 latency (108.7 ms vs 2,668.0 ms), 37× higher throughput (449.3 vs 12.1 pages/min), and 22.2× lower cost per 1,000 pages ($0.048 vs $1.061) on the same RTX 4090 at $0.76/hr (summary_metrics.csv, latency_p50_ms / pages_per_minute / cost_per_1000_pages, sroie_2019 rows).
Why do both engines score so badly on CORD receipts?
Two compounding causes that the benchmark protocol keeps apart from the SROIE ranking: a genuine language mismatch (Indonesian receipts outside both engines’ training focus) and annotation-structure inflation inside CORD’s ground-truth text — CER lands at 0.9101 (docTR) and 0.8959 (Surya2) (summary_metrics.csv, cer, cord_v2 rows). Field metrics absorb part of the shock once an LLM is added (0.5500 vs 0.5203), but CORD rows are quoted and never pooled into any combined ranking.
Which engine should a receipt pipeline pick, docTR or Surya2?
It depends on the axis your pipeline consumes. For raw text at volume with metered cost, docTR’s envelope (108.7 ms, 449.3 pages/min, $0.048/1K pages) dominates. For structured fields without any postprocessor, Surya2’s out-of-box regex F1 (0.3183 vs 0.0766) is the better starting point. For final field quality with LLM postprocessing the choice barely matters (0.6171 vs 0.6139). These results hold for English and Indonesian receipts on one GPU tier in August 2026; any production decision should re-run on the target corpus (see Limitations).
Where do the numbers on this page come from?
Every figure is a row of the first-party benchmark’s published CSVs — results/summary_metrics.csv (CER/WER, field F1, latency, cost, throughput) and results/field_method_comparison.csv (regex vs LLM postprocessing, llm_model = deepseek-v4-flash) — hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run for environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports a head-to-head slice of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims, and not a vendor comparison page. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Both engines saw the same images, same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass, so latency figures are steady-state). Both runs completed with error_rate 0.0 (summary_metrics.csv error_rate column). The underlying run contains eight engines total; this page compares only the two named engines, and the complete 8-engine results are published separately on Traditional OCR vs Document Parsing VLMs.
Runtime Environment
- Hardware: both engines ran on the same NVIDIA RTX 4090 (24 GB); GPU cost computed at the RunPod on-demand rate of $0.76/hr, price timestamped in each run’s redacted manifest (August 2026).
- Engines: out-of-the-box, no fine-tuning. Versions locked: docTR v1.0.1 (traditional two-stage neural OCR, GPU) and Surya2 (surya-ocr 0.22.1) (document-parsing VLM, vLLM-served) — per the public repo model table (README.md) and run manifests.
- LLM postprocessor: deepseek-v4-flash via API at temperature 0 for deterministic output (the llm_model column in field_method_comparison.csv); it was the single model used for all LLM field rows on both engines.
- Cost basis: wall-clock runtime × $0.76/hr, including model initialization — batch processing lowers per-page cost.
- Field postprocessing: SROIE regex field metrics are
postprocessed_sroie_receipt_regex_*(field_method_comparison.csv regex_* columns) — fields extracted from OCR text by a fixed pattern set. They measure OCR + downstream extraction, not native structured output by either model; the LLM_* columns measure OCR text + LLM extraction. The two pipelines are never blended.
Metric Definitions
- CER (Character Error Rate): edit distance (insertions + deletions + substitutions) between OCR text and ground truth, divided by ground-truth characters. Lower is better. Sensitive to case and formatting conventions — mildly conservative against Surya2’s normalized output (case folding, label/value merging).
- WER (Word Error Rate): the same edit-distance calculation at word granularity.
- Field-value F1 (regex): precision/recall harmonic mean over extracted field values using fixed regex patterns on OCR text (traditional OCR + rule-based KIE pipeline). Column: regex_field_value_f1. A score of 0 means no field values recovered.
- Field-value F1 (LLM): the same metric on the LLM postprocessor’s output (OCR text → deepseek-v4-flash → fields). Column: llm_field_value_f1. The two pipelines are different and never blended.
- Document-fields exact: fraction of documents where all target fields matched exactly — a much harsher bar than per-field F1.
- Latency p50/p95 & pages/min: steady-state per-page inference time (warm-then-scored, excludes model loading) and wall-clock throughput including model init.
- Cost per 1,000 pages: billed GPU hours for 1,000 pages at the recorded $0.76/hr rate.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 models × 2 datasets. Columns: model, compute_type, dataset, cer, wer, field_f1_regex, field_acc_regex, latency_p50_ms, latency_p95_ms, cost_per_1000_pages, pages_per_minute, error_rate. Every CER/WER, latency, cost, and throughput number on this page traces to the doctr and surya2 rows here.
- field_method_comparison.csv (GitHub raw). 16 rows; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field-value accuracy and F1, document-fields-exact, llm_median_latency_ms, token counts. Every regex/LLM field-F1 number traces to the doctr and surya2 rows here.
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol, and dataset sample lists (fixed test splits) for reproduction.
- results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with model versions, GPU/driver, torch/CUDA/Python versions, cost metadata with price timestamp, and artifact hashes.
- Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- Document scope — receipts only: SROIE + CORD. Nothing here measures the layout/table/formula/long-document handling where document-parsing VLMs claim their biggest advantages; Surya2’s marketed strengths are unmeasured. Do not use this page to conclude "docTR wins on everything."
- Sample size: 361 English + 100 Indonesian receipts. Field F1 and CER are corpus-sensitive; single-digit differences of a few hundredths (including the 0.006 CER gap and the 0.003 LLM-F1 gap) should be treated as noise, not engineering truth.
- Single GPU tier and single price: all numbers come from one RTX 4090 at $0.76/hr, price timestamped August 2026 in the run manifests. Other GPUs, multi-GPU serving, batch scheduling, or price changes will shift latency, throughput, and cost — re-derive costs at current rates before budgeting.
- Single LLM postprocessor: all LLM rows use deepseek-v4-flash at temperature 0. A different LLM shifts absolute field F1; the convergence ordering may move at the margins. LLM latency (~2.0–2.3 s median, field_method_comparison.csv llm_median_latency_ms) is API-incurred and not part of either engine’s own latency.
- Regex tuning: the pattern set was written once per dataset. A per-format, heavily tuned pattern library could score higher on its own layouts — at the maintenance cost the LLM removes.
- CORD CER is not a per-model quality reading: CORD ground truth embeds annotation structure and neither engine was trained predominantly on Indonesian; CORD CER (0.90–0.91) reflects language mismatch + ground-truth inflation. CORD rows are quoted with framing and never merged into any SROIE ranking (protocol rule).
- CER fairness for VLMs: CER scores exact character matches, so Surya2’s case-folded, label-merged output is mildly penalized for output conventions, not misreads (see related reference on CER). The CER tie therefore slightly understates Surya2; field metrics are the fairer cross-family bar.
- Two engines only: this head-to-head deliberately excludes the other six engines of the underlying run, cloud/API OCR services, and hosted VLM APIs; their latency and pricing models differ fundamentally from the local engines measured here.
- Version pinning: results hold for docTR v1.0.1 and Surya2 0.22.1 (August 2026). Newer releases of either engine may shift every number on this page.
Related references: Traditional OCR vs Document Parsing VLMs · where regex breaks on real documents · why character counts mislead field extraction · Receipt OCR Accuracy
Related reading: AI OCR vs classic OCR accuracy · image data extraction vs OCR engines · AI Document Extraction Pricing (2026)