Regex vs LLM Field Extraction for Receipts
First-Party Benchmark Results (2026)
Last reviewed: 2026-08-14 · Run tier: official · First-party benchmark · 8 models × 2 receipt datasets
What this page does NOT cover: Full OCR text-quality metrics (CER/WER) as the headline — they appear here only as causal context for why one engine's LLM extraction lags. Also not covered: cloud/API OCR engines, fine-tuned document-AI models, OCR engine latency or cost (see Receipt OCR Accuracy and OCR accuracy by document category), or vendor performance claims.
All numbers below come from the benchmark's results/field_method_comparison.csv (regex vs LLM field metrics) and results/summary_metrics.csv (CER/WER context), mirrored in the public GitHub repository and cited row-by-row. F1 values are stored as 0–1 decimals in the CSV and shown as percentages here. The two field-F1 definitions — the regex-based approach and the LLM-based approach — are labeled on every figure and never blended.
When the task is extracting structured fields from OCR text, the choice of postprocessor matters more than the choice of OCR engine. Replacing regex rules with an LLM (deepseek-v4-flash) lifts field F1 from 0.08–0.34 to 0.37–0.62 on English receipts — and from 0.00–0.34 to 0.16–0.55 on Indonesian receipts — across all 8 engines, every engine, on every row of the benchmark.
The inversion to remember: docTR had the worst regex field F1 of all 8 engines (0.077 on SROIE) and the best LLM field F1 (0.617). Its OCR text was fine — regex patterns just could not survive the format variance (dates, currencies, multi-line addresses) of real receipts. The same text that regex mined into 7.7% of fields yielded 61.7% under an LLM. Upstream OCR only needs to be "good enough"; the postprocessor decides how much of that text becomes usable fields.
SROIE Results: English Receipts (361 Samples)
On English receipts, the LLM beats regex for every one of the 8 engines — the smallest lift is 1.76× (PaddleOCR-VL 0.337 → 0.592), the largest 8.1×. And the LLM nearly erases the model gap: 6 of the 7 non-Tesseract engines land inside a 0.05-point band (0.569–0.617), where their regex results were spread across 0.26 points (0.077–0.338).
SROIE (Scanned Receipt OCR and Information Extraction, ICDAR 2019) is the standard English receipt benchmark: 361 test receipts with four target fields — company, date, address, and total (Huang et al., 2019). Each engine first produced OCR text on a shared RTX 4090 setup; that text was then fed to two parallel postprocessors: a fixed set of regex patterns (the traditional OCR + rule-based KIE approach) and the LLM deepseek-v4-flash with a structured extraction prompt. Both scored against the same ground-truth fields. The chart shows regex field F1 (blue) vs LLM field F1 (green) per engine.
Source: field_method_comparison.csv — rows for dataset=sroie_2019, columns regex_field_value_f1 / llm_field_value_f1 (stored as 0–1 decimals, shown as %). LLM postprocessor: deepseek-v4-flash (llm_model column). 361 samples per engine (llm_ok_count).
| OCR Engine | Type | regex F1 | LLM F1 | regex Acc | LLM Acc |
|---|---|---|---|---|---|
| Tesseract | Traditional (CPU) | 23.3% | 43.9% | 21.4% | 43.4% |
| PaddleOCR | Traditional | 32.5% | 58.1% | 29.5% | 58.1% |
| EasyOCR | Traditional | 14.8% | 37.2% | 12.7% | 37.1% |
| docTR | Traditional | 7.7% | 61.7% | 6.2% | 61.7% |
| Docling | Pipeline parser | 22.4% | 56.9% | 20.4% | 56.6% |
| Surya2 | Document VLM | 31.8% | 61.4% | 30.0% | 61.4% |
| Unlimited-OCR | Document VLM | 33.8% | 60.5% | 30.9% | 60.5% |
| PaddleOCR-VL | Document VLM | 33.7% | 59.2% | 31.0% | 58.5% |
Source: field_method_comparison.csv — sroie_2019 rows: regex_field_value_f1 / llm_field_value_f1 / regex_field_value_accuracy / llm_field_value_accuracy (0–1 decimals shown as %). docTR F1 0.0766 → 0.6171; Surya2 0.3183 → 0.6139; PaddleOCR 0.3254 → 0.5810.
CORD Results: Indonesian Receipts (100 Samples, Stress Test)
CORD widens the gap, not because the LLM gets better — it does not — but because regex collapses: six of eight engines score regex field F1 at or below 10.8%, and two (docTR 0.0%, EasyOCR 0.7%) recover almost nothing. The LLM still lifts every non-Tesseract engine to 33.8% or higher, proving its extraction generalizes across languages and layouts where hand-written patterns cannot.
CORD v2 is an Indonesian-language receipt dataset with nested fields (menu, sub_total, total) (Park et al., 2019), run as a cross-language stress test: none of the 8 engines was trained primarily on Indonesian receipts. Two caveats apply to reading this table. First, CORD's ground-truth text includes annotation structure and VLM output-normalization differences, which systematically inflate raw CER for every engine — so the field metrics below, not CER, are the fair cross-model comparison. Second, the regex patterns were written for the flat English SROIE schema; CORD's nested schema and Indonesian formatting (Rp currency, date conventions) defeat them — which is itself the finding: rules tuned to one market do not travel.
Source: field_method_comparison.csv — rows for dataset=cord_v2, columns regex_field_value_f1 / llm_field_value_f1 (0–1 decimals shown as %). 100 samples per engine (llm_ok_count).
| OCR Engine | Type | regex F1 | LLM F1 | regex Acc | LLM Acc |
|---|---|---|---|---|---|
| Tesseract | Traditional (CPU) | 7.5% | 16.3% | 5.6% | 14.3% |
| PaddleOCR | Traditional | 1.5% | 55.3% | 1.1% | 50.7% |
| EasyOCR | Traditional | 0.7% | 33.8% | 0.4% | 30.5% |
| docTR | Traditional | 0.0% | 55.0% | 0.0% | 51.8% |
| Docling | Pipeline parser | 6.1% | 46.9% | 4.6% | 44.3% |
| Surya2 | Document VLM | 24.6% | 52.0% | 18.9% | 49.0% |
| Unlimited-OCR | Document VLM | 10.8% | 46.8% | 8.4% | 45.0% |
| PaddleOCR-VL | Document VLM | 34.1% | 52.0% | 28.7% | 49.3% |
Source: field_method_comparison.csv — cord_v2 rows (regex/llm field-value F1 and accuracy, 0–1 decimals shown as %). PaddleOCR LLM F1 0.0154 → 0.5527; docTR 0.0 → 0.5500; tesseract 0.0752 → 0.1627.
Why the LLM Flattens the Gap — and Where Its Ceiling Is
Three patterns in the numbers explain what is happening under the surface. Each is a testable claim with its exact CSV cells below.
(a) The LLM turns a 0.26-point model gap into a 0.05-point band
Under regex, which engine you picked mattered enormously: SROIE field F1 spanned 0.077 (docTR) to 0.338 (Unlimited-OCR) — a 0.26-point spread. Under the LLM, six of the seven non-Tesseract engines land between 0.569 (Docling) and 0.617 (docTR) — a 0.05-point band (field_method_comparison.csv, sroie_2019 rows, regex_field_value_f1 vs llm_field_value_f1). The upstream OCR only needs to produce readable text; the LLM extracts fields from it with roughly engine-independent quality. Two engines trail this band: EasyOCR at 0.372 (the lowest LLM F1 on English receipts) and Tesseract at 0.439. On CORD, Tesseract becomes the clear laggard — its 0.163 is the worst LLM result in the entire table.
(b) Tesseract's ceiling is set by its OCR, not its postprocessor
"Garbage in, garbage out" applies even to LLM postprocessing. Tesseract's raw OCR text on Indonesian receipts has a character error rate of 0.9523 (summary_metrics.csv, tesseract/cord_v2, cer) — roughly 95 characters in 100 are wrong or misordered. Its LLM field F1 on CORD is consequently 0.1627 (field_method_comparison.csv, tesseract/cord_v2, llm_field_value_f1): an LLM cannot extract a company name or total from text it cannot read. The ceiling of any postprocessor is set by the OCR base quality beneath it — a constraint no prompt engineering removes.
(c) docTR is the biggest beneficiary: regex-floor to LLM-top
docTR's story is the cleanest demonstration that the problem was the postprocessor, not the OCR. Its regex field F1 on SROIE was the lowest of all 8 engines at 0.0766 — its clean, accurate text (SROIE CER 0.1971, the best in the benchmark, summary_metrics.csv row for doctr/sroie_2019) simply did not match the regex patterns for totals with thousands separators or multi-line addresses. Under the LLM the same text yields the benchmark's highest SROIE result, 0.6171 — an 8.1× lift and the largest in the table (field_method_comparison.csv, doctr/sroie_2019, regex_field_value_f1 0.0766 → llm_field_value_f1 0.6171). On CORD the effect repeats: 0.0 under regex, 0.5500 under the LLM.
When Regex Is Enough vs When to Use an LLM
The honest tradeoff is not "regex is broken" but "regex is brittle where receipts are variable." Regex postprocessing is deterministic, free, and instant; an LLM postprocessor adds tokens and 1.8–2.4 s median per document (field_method_comparison.csv, llm_median_latency_ms across all 16 rows). In return, the LLM recovered 1.76–8.1× more fields on SROIE — and on CORD, where regex collapsed to near zero for most engines, the LLM recovered 33.8–55.3% of fields that rules-based extraction simply could not reach.
When regex is enough: if your documents come from a small, stable set of layouts and the fields you need appear in near-constant formats — a fixed invoice-number pattern, a single date convention, one currency — regex is the right tool: it costs nothing, runs in microseconds, and its failures are predictable. The benchmark's floor cases show this: engines with receipt-aligned regex still reached 0.34 field F1 on SROIE (Unlimited-OCR 0.3376, PaddleOCR-VL 0.3368).
When to use an LLM: as soon as format variance enters — multiple currencies, regional date formats, multi-line addresses, vendor names in varying styles, or a second language. Each of those breaks a pattern; the LLM absorbs them all in one prompt. The benchmark's CORD results quantify what "one more language" costs a rules-based pipeline: regex field F1 fell from 0.08–0.34 (English SROIE) to 0.00–0.34 (Indonesian CORD), while the LLM held 0.16–0.55 — a penalty the regex approach paid 100% of and the LLM paid only partially. The right architecture for most production flows is hybrid: LLM extraction for variable fields, regex or rule-based validation for fields with strict expected formats, with the 1.8–2.4 s per-document LLM cost amortized in async batch processing rather than synchronous user waits.
Frequently Asked Questions
Is LLM field extraction more accurate than regex extraction?
Yes, in this benchmark, on every row: LLM postprocessing (deepseek-v4-flash) beat regex postprocessing for all 8 OCR engines on both datasets (field_method_comparison.csv, all 16 rows). On SROIE English receipts, LLM field F1 was 0.37–0.62 vs regex 0.08–0.34; on CORD Indonesian receipts, 0.16–0.55 vs 0.00–0.34.
Which OCR model extracts receipt fields best with LLM postprocessing?
docTR on SROIE, at 0.6171 field F1 — but the differences among engines mostly disappear once an LLM is in the loop: six of seven non-Tesseract engines land between 0.569 and 0.617 (field_method_comparison.csv, sroie_2019 rows, llm_field_value_f1). On CORD, PaddleOCR leads at 0.5527, with docTR a close second at 0.5500.
Why does Tesseract fall behind even with an LLM?
Because its OCR text is too degraded for any postprocessor to recover fields from. On Indonesian receipts its character error rate is 0.9523 (summary_metrics.csv, tesseract/cord_v2), which caps its LLM field F1 at 0.1627 — the worst LLM result in the benchmark (field_method_comparison.csv, tesseract/cord_v2, llm_field_value_f1).
How much does an LLM improve field extraction over regex?
Between 1.76× and 8.1× on English receipts depending on the engine, with the largest lift on docTR (0.077 → 0.617 field F1, field_method_comparison.csv sroie_2019 rows). On Indonesian receipts the lift is far larger — 36× for PaddleOCR and effectively unbounded for docTR (0.0 → 0.55) because regex recovered almost nothing.
Does LLM postprocessing get every field right?
No — and readers should not expect it to. The best LLM field F1 in the benchmark is 0.617 (docTR, SROIE), meaning roughly 38% of fields were still missed or wrong, and the best whole-document exact-match rate is 15.5% (Surya2, SROIE, llm_document_fields_exact) — i.e., at most ~1 in 6 documents had all four fields exactly right. LLM extraction is a large improvement over regex, not a silver bullet; field-level confidence scoring and human review remain necessary for production.
How slow or expensive is LLM postprocessing?
In this benchmark the LLM call added 1.8–2.4 s median latency per document (field_method_comparison.csv, llm_median_latency_ms, all 16 rows), and the full 16-group run consumed roughly 1.33M prompt + 0.28M completion tokens (sum of llm_prompt_tokens / llm_completion_tokens). This is not a real-time field-extraction latency — it suits async batch processing, not interactive lookups.
Does LLM postprocessing work on non-English receipts?
Better than regex, but with a smaller ceiling. On Indonesian CORD receipts the LLM held field F1 at 0.16–0.55 where regex collapsed to 0.00–0.34 (field_method_comparison.csv, cord_v2 rows) — but the best CORD result (0.5527) still trails the best English result (0.6171), reflecting both the harder script and the degraded OCR base. Rules written for English receipts essentially stopped working; the LLM degraded gracefully instead.
Methodology & Sources
Protocol
This page reports the field-extraction comparison of the ImageToTable.ai open-source OCR benchmark — an independent, reproducible experimental run (official tier), not a survey of third-party claims. The pipeline for each (model × dataset) pair: the OCR engine produces text from receipt images; that text is then extracted twice — once by a fixed set of regex patterns and once by the LLM postprocessor — and both outputs are scored against the same ground-truth fields. The LLM postprocessor is deepseek-v4-flash (the llm_model column in the comparison CSV), run at temperature 0 for deterministic output. All 361 SROIE and 100 CORD samples completed successfully (llm_ok_count = 361 / 100).
Runtime Environment
- Hardware: all GPU runs on NVIDIA RTX 4090 (24 GB). The PyTorch-based engines (docTR, EasyOCR, Docling) ran with PyTorch 2.8.0+cu128 (CUDA 12.8); Tesseract ran CPU-only (no GPU cost); the vLLM-served engines (Surya2, Unlimited-OCR, PaddleOCR-VL) ran on a vLLM pod. Driver and Python versions vary slightly per run and are recorded exactly in the redacted run manifests (see Artifact Access below).
- LLM postprocessor: deepseek-v4-flash via API, temperature 0.
- Measurement mode: warm_then_scored — a fixed warm-up pass precedes the scored pass, so latency figures are steady-state.
- Datasets: SROIE 2019 test — 361 English receipts, flat fields (company, date, address, total), CC-BY-4.0; CORD v2 test — 100 Indonesian receipts, nested fields (menu, sub_total, total), CC-BY-4.0.
- Sample counts: SROIE 361 / CORD 100, all from the fixed test splits (no training data leakage).
The upstream OCR text consumed by both postprocessors came from these engine versions:
| OCR Engine | Version | Backend |
|---|---|---|
| Tesseract | 5.3.4 | CPU (no GPU) |
| PaddleOCR | 3.7.0 | PaddlePaddle-GPU 3.3.1 |
| EasyOCR | 1.7.2 | PyTorch |
| docTR | 1.0.1 | PyTorch |
| Docling | 2.119.0 | PyTorch |
| Surya2 | 0.22.1 | vLLM |
| Unlimited-OCR | baidu/Unlimited-OCR | vLLM |
| PaddleOCR-VL | 1.6 | vLLM |
Full per-run reproducibility fingerprints are in the benchmark's redacted run manifests under results/manifests/ in the public repository (one manifest.json per published run, 16 total). Each discloses: run id, model + version, runner-script hash (SHA-256), GPU model/driver/VRAM, torch/CUDA/torchvision/torchaudio versions, Python version, pip-freeze hash, cost metadata (GPU $/hour and price timestamp), measurement mode, and artifact hashes (sample list, predictions, metrics, performance).
Metric Definitions
- Field-value F1 (regex approach): harmonic mean of precision and recall over extracted field values, using regex postprocessing — the traditional OCR + rule-based KIE approach. Column: regex_field_value_f1.
- Field-value F1 (LLM approach): the same metric computed on the LLM postprocessor's output — the OCR + LLM postprocessing approach. Column: llm_field_value_f1. These two approaches are different pipelines and are never blended.
- Field-value accuracy: fraction of extracted field values that exactly match ground truth (regex_field_value_accuracy / llm_field_value_accuracy).
- Document-fields-exact: fraction of documents where every target field matched exactly (regex_document_fields_exact / llm_document_fields_exact).
- CER/WER (context only): character/word error rate of raw OCR text, used on this page only to explain Tesseract's LLM ceiling.
Artifact Access
- field_method_comparison.csv (GitHub raw). 16 rows = 8 models × 2 datasets; columns model, dataset, llm_model (= deepseek-v4-flash), regex/llm field F1 + accuracy, document-fields-exact, llm_median_latency_ms, llm_prompt_tokens, llm_completion_tokens. Every field-F1 and accuracy number on this page traces to a row here.
- summary_metrics.csv (GitHub raw). CER/WER per model × dataset, used for the Tesseract-ceiling explanation (tesseract/cord_v2 cer 0.9523) and docTR's SROIE CER 0.1971.
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, run manifests, and dataset sample lists for reproduction.
- results/manifests/ (GitHub). One redacted
manifest.jsonper published run (16 runs), with the per-run environment fingerprint and artifact hashes listed above. - Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE dataset definition and license.
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2019). CORD v2 dataset definition and license.
Limitations
- Document scope: Receipts only (SROIE + CORD). Findings do not generalize to invoices, forms, or long documents without further testing.
- Sample size: 361 English + 100 Indonesian receipts. Field F1 is sensitive to corpus composition; treat single-point differences of a few hundredths as noise.
- Single LLM model: All LLM rows use deepseek-v4-flash. A different LLM (size, prompting, or vendor) would produce different absolute numbers; the regex-vs-LLM ordering may shift at the margins.
- CORD ground-truth caveat: CORD gt_text includes annotation structure and VLM normalization differences, so raw CER on CORD is systematically inflated for every engine (e.g., PaddleOCR-VL CER 1.08 is an artifact, not a real reading of its text quality). Field metrics are the fair comparison; CER is used here only for the Tesseract explanation.
- Latency is not real-time extraction latency: llm_median_latency_ms (1.8–2.4 s) covers the whole OCR-text → LLM field extraction call, not per-field lookup, and does not include OCR time itself.
- Regex tuning: The regex patterns are a fixed set written once per dataset (SROIE-flat schema); a heavily tuned per-vendor regex library could score higher on its own formats — at the cost of the maintenance burden the LLM removes.
Related references: the gap between field and character accuracy · Receipt OCR Accuracy · document-type accuracy benchmarks
Related reading: How to Read OCR Accuracy Claims