References
Data references and terminology definitions for document AI. Each page is source-verified and designed to be cited.
Statistics
First-party head-to-head: docTR vs Docling on 361 SROIE + 100 CORD receipts — 6.7x latency, 8.3x cost, 3.0x CER, 2.9x regex-field flip. Methodology.
First-party head-to-head: docTR vs Surya2 on 361 SROIE + 100 CORD receipts — CER tie, 4.2x field-F1 flip, 22x cost gap, methodology.
First-party benchmark: EasyOCR vs docTR on 361 SROIE + 100 CORD receipts — docTR 30% lower CER, 1.66x LLM field F1; EasyOCR 1.93x regex F1.
First-party benchmark: OCR cost per 1,000 pages for 8 open-source engines on RTX 4090 — $0.048 (docTR) to $1.061 (Surya2). SROIE + CORD, methodology.
Per-page latency for 8 open-source OCR engines on RTX 4090 — 108.7 ms (docTR) to 2,668 ms (Surya2) p50. p95 tails, throughput, methodology.
First-party head-to-head: PaddleOCR vs EasyOCR on receipts — CER 0.204 vs 0.283, 2.2x regex F1, 2x cost gap, LLM paradox. Methodology.
First-party benchmark: Surya2 vs Unlimited-OCR vs PaddleOCR-VL on 361 SROIE + 100 CORD receipts — field F1, CER fairness, cost, latency.
First-party head-to-head: Tesseract (CPU) vs PaddleOCR (GPU) on 361 SROIE + 100 CORD receipts — CER, field F1, throughput parity, cost. Methodology.
First-party benchmark: 8 open-source OCR engines vs document-parsing VLMs on 361 SROIE + 100 CORD receipts — CER, field F1, latency, cost. Methodology included.
Why VLM CER is inflated: case folding and GT structure add phantom errors. First-party SROIE + CORD data — field F1 is the fair cross-family metric.
First-party benchmark: LLM postprocessing beats regex for receipt field extraction, field F1 for 8 OCR models on SROIE + CORD, methodology included.
First-party benchmark of 5 open-source OCR models on 461 receipts (SROIE + CORD): CER, WER, field F1, latency, cost per 1,000 pages.
Exception rates in document automation: 9% best-in-class vs 22% average, root causes, review costs, and threshold tradeoffs with all sources cited.
Independent benchmarks on document processing cost per record: manual $8-16, template $2-5, AI from $0.02/page. By document type, size, geography.
Handwriting OCR accuracy data: IAM CER 1.2-1.7%, Arabic 8.9-16%, historical 71-82%. Broken down by script, model & document type. 15 citable sources.
Benchmarks on invoice processing time: 12.5 min manual touch vs 9.2-day cycle, by method, complexity, stage — with full sources for citation.
Independent benchmarks on data entry automation ROI: 3-12 month payback, 60-80% savings, breakeven volume, plus a formula to model your return.
Independent benchmarks on OCR accuracy by document type: 99.2-99.8% digital PDFs, 98-99% scans, 46-95% handwriting, table TEDS. 15 citable sources.
Receipt OCR accuracy by condition and field: 90%+ clean scans, 55-73% faded thermal, 67-69% phone photos. Totals 95-99%, merchants 73-88%. 14 sources.
Independent benchmarks on manual data entry error rates: field-level 1-4%, record-level up to 40%, by document type, industry, and role.
Definitions
Data capture turns paper, scans, PDFs, and photos into machine-readable data. How it works, its types, and misconceptions.
ICR is machine-learning character recognition for hand-printed text. How it evolved from OCR, the ICR vs OCR vs IWR differences, and where it's used.
Key Information Extraction (KIE) pulls fields like invoice numbers and dates from documents into structured data. How it works, boundaries, misconceptions.
A searchable PDF is a scan with an invisible OCR text layer — you can select, copy, and Ctrl+F the text. How it works, PDF/A rules, and common traps.
Field-level accuracy counts correctly extracted fields; character-level counts correctly read characters. 99% character accuracy is not 99% accurate data.
IDP is AI software that captures, classifies, extracts, and validates data from documents. How it works, evolution from OCR to LLMs, and misconceptions.
STP rate measures the share of documents processed with zero human touch — industry avg 32%, best-in-class 49%, AI 60-80%. Formula and misconceptions.
OCR converts images of printed, handwritten, or typed text into machine-readable data. How it works, the 4 types of OCR, its history, and use cases.