OCR Accuracy by Document Type: Digital PDF, Scan,Handwriting and Table Benchmarks (2026)

Last reviewed: 2026-08-10 · Data coverage: Partial · Sources: 16 independent studies

What this page covers: Character-, word-, field-, and structure-level OCR accuracy benchmarks aggregated from peer-reviewed studies (Springer, UNLV ISRI, ICDAR competitions), government bodies (NIST, EDPB, US GPO), and methodology-transparent industry benchmarks (AIMultiple). Broken down by document type, scan resolution, field type, and language — with the character-vs-field-vs-structure metric distinction that most "99% accuracy" claims omit. All sources are third-party and independently verifiable.
What this page does NOT cover: Manual data entry error rates (see Manual Data Entry Error Rates), the mechanism of OCR itself (see What is OCR?), or vendor-specific claims about any single product beyond the benchmark figures cited below.

All numbers below are based on publicly available third-party studies and industry reports cited in the methodology section. Where sources disagree, conservative ranges (lowest–highest reported values) are used. This page is a directional benchmark, not a precise forecast for any single engine or organization.

Every OCR vendor quotes a version of "99% accuracy". Across 16 independent studies, that number is true for exactly one document condition — clean digital-born text, where all major engines score above 99.2% — and it becomes progressively less true for everything else: clean scans drop to 98–99%, historic and degraded scans to 71–98%, handwriting to a 46–95% spread across solutions, and damaged originals to as low as 72.7%. The figure is not a lie; it is a document-type-specific measurement presented as if it were universal.

99.2–99.8%
Character accuracy on clean digital-born / typed text — all engines above 99.2% on easy typed text (IntuitionLabs 2025); ABBYY reports up to 99.8% on clean scans (vendor-claimed)
71–98%
Raw character accuracy across degraded and historic newspaper scans (1803–1954), measured by the National Library of Australia (Holley 2009) — the widest band of any common document class
46–95%
Handwriting recognition accuracy across 12 modern solutions (AIMultiple 2026) — top LLM-based systems hit 93–95%, legacy engines sit near the bottom of the band

Where the "99%" Figure Comes From — and Why It Misleads

The 99% figure is a character-level measurement on clean printed text, and character-level accuracy is the most favorable lens available. Understanding that — and the three other lenses in use — explains why "99% OCR accuracy" and "78% field accuracy" can both be true for the same system.

Character Error Rate (CER) counts individual misread characters: a CER of 1% means roughly one wrong character per hundred. The widely adopted benchmark scale for printed text, formalized by the National Library of Australia's newspaper digitization program, defines good OCR as 98–99% accurate (1–2% CER), average as 90–98% (2–10% CER), and poor as below 90% (Holley, 2009). The European Data Protection Board's OCR guidance independently puts common printed-document performance at 95–99% character accuracy, with 98–99% treated as the production-grade guarantee level (EDPB, 2024).

The problem is that character accuracy is the least demanding of the metrics in use. A document with 2,000 characters and a 99% character accuracy still contains ~20 wrong characters. At the word level, that same output might score 96–98% because a single misread character fails the whole word. At the field level — the measure that matters for extracting an invoice number or a date — one wrong digit fails the entire field, so a system with 99% character accuracy can deliver only 78–98% field accuracy depending on the engine (BusinessWareTech benchmark, 2025). For tables, a fourth metric applies: Tree Edit Distance Similarity (TEDS), which scores whether the extracted table structure and cell content match the original — an entirely different axis from character accuracy.

The historical record shows how hard the clean-text ceiling is. UNLV's Information Science Research Institute ran annual OCR accuracy tests from 1992–1995 on machine-printed English pages; best commercial systems climbed from roughly 94% to 98%+ character accuracy and largely plateaued there (Rice, Jenkins & Nartker, 1995). NIST reached the same conclusion in its 1998 evaluation: a 1% character error rate is attainable only when the source is a clean, fixed-format typed original (NIST IR 6101, 1998). Twenty-five years of deep learning have pushed that ceiling, but only for the cleanest inputs.

Document type is the single largest driver of OCR accuracy — larger than the choice of engine. The chart below plots the published range for each document class; the spread between the lowest and highest observed value is the honest answer, not a single "average." Degraded or handwritten documents produce ranges so wide that a point estimate would be actively misleading.

Breakdown by Document Type

OCR accuracy ranges by document type: digital-born text 99.2–99.8%, clean scanned print (300 DPI) 98–99%, degraded/historic scans 71–98%, handwriting best modern systems 94–99%, handwriting full solution range 46–95%, damaged originals 72.7–94.7%.

Sources: IntuitionLabs (2025) — all engines >99.2% on easy typed text; Holley (2009) — clean 98–99% "good" benchmark; degraded 71–98.02% raw character confidence on Australian newspapers 1803–1954; AIMultiple (2026) — handwriting 46–95% across 12 solutions; GatedLexiconNet (2024) — IAM word accuracy ~94.3%; IJSAT (2026) — damaged-original 72.7–94.7% at 300 DPI. Conservative ranges used where sources disagree.

Document Type / ConditionAccuracy RangeMidpointSourceYearSample
Digital-born PDF (ERP/typed output)99.2–99.8%99.5%IntuitionLabs; ABBYY (vendor claim)2025All engines >99.2% on easy typed text
Clean scanned print (300 DPI)98–99%98.5%Holley; EDPB2009 / 2024"Good OCR" benchmark; NLA digitized collection
Scanned — degraded or historic71–98%~85%Holley200930+ newspaper titles, 1803–1954, raw character confidence
Handwriting — best modern systems (benchmarks)94–99% word96%GatedLexiconNet; LLM HTR benchmark2024–2025IAM (657 writers, 13,353 lines); RIMES (French letters)
Handwriting — cross-solution range46–95%~70%AIMultiple202612 solutions, SBERT similarity scoring
Photo of damaged original72.7–94.7%~84%IJSAT2026EasyOCR; normal vs crumpled/dirty/wet, 300 DPI
Tables — structure (born-digital)TEDS 0.91–0.970.94PubTabNet leaderboard (MuTabNet 0.969); TrOCR-ctx (0.909)2024–2025PubTabNet corpus, TEDS structural similarity
Tables — detection (archival / hand-drawn)wF1 0.49–0.530.51ICDAR 2019 cTDaR2019Archival track, weighted F1 at IoU 0.6–0.9

Sources listed in table above. Character/word figures come from peer-reviewed and government studies; vendor-claimed figures are labeled as such. TEDS and wF1 are structure-level metrics and are not directly comparable to character accuracy — see Why Accuracy Ranges Are So Wide. For the widest-degradation row, the 71% floor is a 2009 measurement on the worst-performing historic newspaper title (the most recent available published raw figure); modern engines on similar documents score materially better (see Methodology).

Resolution is the most controllable accuracy lever, but the published evidence is thinner than vendor blogs suggest. The one rigorous government study — the US Government Publishing Office's test on older and discolored documents — found that color (RGB) capture, not higher DPI, was the decisive factor: RGB scanning met the GPO's 99% character accuracy requirement, while bitonal (black-and-white) scans of the same documents scored ~77%, and downsampling from 300 to 200 DPI produced no accuracy gain (US GPO white paper).

Breakdown by DPI and Scan Quality

Scan Resolution / ModeObserved EffectSourceYear
300 DPI+ (standard)Meets 98–99% character accuracy on clean print; 400–600 DPI recommended for fonts under 10ptUniv. of Pittsburgh OCR guide; US GPO2024 / ~2022
200 DPIAdequate on clean print; downsampling from 300 did not change recognition in GPO testingUS GPO~2022
Below 150 DPIRated a high-impact degradation factor; small fonts fail and accuracy falls below the 98–99% clean bandLlamaIndex glossary; degradation floor per IJSAT (2026)2026
72 DPI (screen capture)Insufficient pixel density for reliable character discrimination; classified "high impact" on accuracyLlamaIndex glossary2026
Color (RGB) vs bitonal scan of aged documentsRGB reached 99% vs ~77% for bitonal on older/discolored originalsUS GPO~2022

Sources listed in table above. Note the honest caveat: no peer-reviewed study isolates DPI alone while holding all other factors constant. The "below 150 DPI" row therefore combines explicit low-DPI guidance from the LlamaIndex reference glossary with the physical-damage accuracy floor measured by IJSAT (2026) at 300 DPI — a proxy, stated as such.

Field type determines whether a reading error matters. A misread word in a paragraph is invisible to most workflows; a misread digit in an invoice number fails the entire field. The only published benchmark that isolates field types across multiple engines is a 2025 invoice-extraction study, and its spread — 40–98% depending on field type and engine — is the clearest available evidence that "OCR accuracy" is not one number (BusinessWareTech, 2025).

Breakdown by Field Type

Field TypeAccuracy RangeContextSourceYear
Key-value text fields (invoice headers)78–98% by engineGPT-4o + OCR 98%, Azure Document Intelligence 93%, Google Document AI 82%, AWS Textract 78%BusinessWareTech2025
Line items / table rows40–82% by engineLine-item detection: Textract 82% vs Google Document AI 40% in the same testBusinessWareTech2025
Numeric fields (dates, totals, IDs)78–98% (within field band)Constrained format helps, but a single misread digit fails the fieldBusinessWareTech; AIMultiple invoice benchmark2025–2026
Checkboxes / mark-senseNo reliable public benchmarkData gap — flagged in Limitations; legacy mark-sense literature predates AI OCR
Signatures (recognition)No public accuracy benchmarkData gap — vendors report detection, not recognition accuracy

Sources listed in table above. The 78–98% and 40–82% figures come from a single independent benchmark (BusinessWareTech, 2025, real-invoice test set) — a C-grade source, corroborated in direction by AIMultiple's invoice benchmark, which found accuracy declines sharply on lower-quality documents across all tools (AIMultiple, 2026). Rows marked "Data gap" have no citable third-party measurement; they are stated rather than filled with vendor numbers.

Language is the least-publicized accuracy dimension because almost no vendor publishes per-language benchmarks. The one peer-reviewed multi-language comparison — a 2021 study of 18,568 English and Arabic document images — found English accuracy "considerably higher" than Arabic on the same engines, with the gap on the order of 10–25 percentage points for out-of-the-box recognition (Hebert et al., 2021). Handwriting benchmarks add a second data point: French (RIMES) and English (IAM) handwriting now reach near-print quality in the best systems, but only on clean, standardized corpus handwriting.

Breakdown by Language and Script

Language / ScriptAccuracyEvidenceSourceYear
English — printed (clean)96–99.5%Baseline for all major engines on clean textHebert et al.; EDPB2021 / 2024
Arabic — printedRoughly 10–25 pts below EnglishSame engines, single-common-font corpus; still considerably lowerHebert et al.2021
English handwriting (IAM benchmark)~94.3% word accuracy SOTAWER 5.73% (GatedLexiconNet); 3.34% WER with GPT-4o-miniGatedLexiconNet; LLM HTR benchmark2024–2025
French handwriting (RIMES benchmark)~97–99% character accuracyCER 0.9–1.7% on business lettersGatedLexiconNet2024
CJK, Devanagari and other scriptsNo comparable public benchmarkData gap — no peer-reviewed cross-engine comparison found

Sources listed in table above. IAM and RIMES are standardized handwriting corpora — real-world messy handwriting, particularly mixed print-cursive, scores below these clean-corpus figures (the 46–95% cross-solution band in the document-type table is the real-world reference). The Arabic gap is a conservative interpretation of the published finding "accuracy for English was considerably higher than for Arabic" on a single-common-font corpus; a harder corpus would likely widen it.

Why Accuracy Ranges Are So Wide: Four Metrics, One Word

The widest apparent contradictions in OCR accuracy data are not measurement errors — they are different metrics applied to different document conditions. "99% character accuracy" and "78% field accuracy" describe the same system; "wF1 0.49 table detection" and "TEDS 0.96 table structure" describe the same technology on different corpora.

The four metrics in use form a strict hierarchy of difficulty. Character accuracy (99%+) is the easiest to inflate. Word accuracy is lower because one bad character fails a word — the Australian newspaper correction study measured 81.7% word accuracy on raw historical output, rising to 97.6% after manual correction (Cassidy, 2019, Macquarie University study of the same corpus class). Field accuracy is lower still because one wrong digit fails the whole field — the 78–98% engine range above. Structural metrics (TEDS for tables, wF1 for detection) are a separate axis entirely: a table can be read with perfect characters and still score poorly if the row/column structure is reconstructed wrong, which is why archival table detection tops out at wF1 0.49–0.53 even as modern table structure recognition reaches TEDS 0.91–0.97 (TrOCR-ctx, 2025; PubTabNet leaderboard).

Noise compounds the spread. In the peer-reviewed Hebert et al. study, adding a single artificial noise layer to clean book scans dropped word accuracy from roughly 96% to 82.5% for the best engine, and to 58.4% for open-source Tesseract; two noise layers pushed further down (Hebert et al., 2021). This is why the honest answer to "how accurate is OCR?" is always a range conditioned on document type — and why any single-number claim should be treated as a clean-text measurement wearing a universal costume.

How to Use This Data

You don't need a full audit to turn "we've heard OCR is 99% accurate" — the claim every vendor repeats, and the one this page tests — into a defensible expectation for your own documents. A three-step back-of-the-envelope calculation using the ranges above is enough, and it usually reveals that the headline figure does not apply to the documents you actually process.

  1. Classify your document mix. Sort your volume into the bands above — e.g., for a typical accounts-payable flow of 1,000 invoices/month: 70% digital-born PDFs (99.2–99.8%), 20% clean scans (98–99%), 10% degraded scans or files with handwritten annotations (~85% midpoint).
  2. Convert character accuracy to field accuracy. Character-level figures flatter extraction workflows. Use the measured field-level band for your engine tier instead: 93% for a mid-tier cloud processor, 98% for an LLM-assisted pipeline, 78–82% for lower-tier engines (BusinessWareTech 2025).
  3. Estimate monthly error count. At 93% field accuracy with 12 fields per invoice, 1,000 invoices yield roughly 840 wrong fields per month (1,000 × 12 × 7%) before any review. If 10% of your volume is degraded or handwritten, that error count rises by the band gap — typically 15–30 percentage points of field accuracy.
  4. Decide whether the count is acceptable. If 840 wrong fields/month exceeds what your downstream systems absorb, the published levers are: raise scan quality toward the 300-DPI/RGB band (GPO found this alone closes the 77%→99% gap on aged documents), route low-confidence fields to human review, or apply the correction stage that lifted historical word accuracy from 81.7% to 97.6% (Cassidy, 2019).

These are rough-estimate formulas, not a substitute for benchmarking your own documents. Every published range on this page was measured on someone else's corpus — the only accuracy number that applies to your workflow is the one you measure on it. The value of this page is setting the expectation before you measure.

Frequently Asked Questions

What is the average OCR accuracy rate?

The average depends entirely on document type. Across the sources on this page, clean digital-born text averages 99.2–99.8% character accuracy (IntuitionLabs 2025; ABBYY), clean scans 98–99% (Holley 2009; EDPB 2024), degraded or historic scans 71–98% (Holley 2009), and handwriting 46–95% across solutions (AIMultiple 2026). A single "average OCR accuracy" number is not meaningful — the range is the answer.

What is a good OCR accuracy rate?

The National Library of Australia's widely adopted scale defines good printed-text OCR as 98–99% character accuracy (1–2% CER), average as 90–98%, and poor as below 90% (Holley, 2009). For field-level extraction — the metric that matters for business data — a more realistic "good" is 93–98%, which only the top-tier engines in the 2025 invoice benchmark reached.

How accurate is OCR on scanned documents vs digital PDFs?

Clean scans of printed text run 98–99% character accuracy — close to the 99.2–99.8% of digital-born text, but not equal to it. The gap widens sharply once scans degrade: historic newspaper scans measured 71–98% raw accuracy (Holley 2009), and damaged originals fell to 72.7–94.7% even at 300 DPI (IJSAT 2026). Low DPI, bitonal capture, and aging paper each push scans down the band.

How accurate is handwriting OCR?

State-of-the-art systems reach ~94–99% word accuracy on clean benchmark corpora like IAM and RIMES (GatedLexiconNet 2024; LLM HTR benchmark 2025), but across all solutions the range is 46–95% (AIMultiple 2026) — legacy engines sit near the bottom, and messy real-world handwriting, mixed print-cursive, or historical scripts score well below the clean-corpus figures.

What DPI is best for OCR accuracy?

300 DPI is the standard recommendation — it meets the 98–99% character accuracy band on clean print, with 400–600 DPI advised for fonts under 10pt (University of Pittsburgh OCR guide). Notably, the GPO's study found scan color mode mattered more than resolution for aged documents: RGB capture hit 99% while bitonal scored ~77%, and downsampling from 300 to 200 DPI changed nothing.

Why is OCR accuracy lower on tables?

Tables add a structure-recognition step on top of character recognition: the engine must reconstruct rows, columns, and merged cells before the content can be used. Modern systems score TEDS 0.91–0.97 on born-digital table corpora (PubTabNet leaderboard; TrOCR-ctx 2025), but archival or hand-drawn tables drop to wF1 0.49–0.53 detection (ICDAR 2019 cTDaR), and line-item extraction in real invoice tests ranged 40–82% by engine (BusinessWareTech 2025).

Does OCR accuracy differ by language?

Yes — materially. The only peer-reviewed cross-language comparison found English accuracy 10–25 percentage points higher than Arabic on the same engines and corpus (Hebert et al. 2021). Handwriting benchmarks show French and English corpora now reach near-print quality in top systems, but no comparable public benchmark exists for CJK, Devanagari, or other scripts — a genuine data gap (see Limitations).

Methodology & Sources

How This Page Was Built

This page aggregates data from 16 independent studies spanning 1995–2026. Sources were selected based on three criteria: (1) publicly documented methodology with stated sample size, (2) data collected within the last 3 years where available, with older sources labeled by year in-text, and (3) relevance to document processing rather than marketing claims. Where multiple sources report the same dimension, conservative ranges (lowest–highest reported values) are used — we prefer to under-claim. Where sources disagree significantly, the disagreement is explained by metric level or corpus rather than averaged away.

The metric distinction is the spine of this page: character-level, word-level, field-level, and structure-level (TEDS/wF1) figures are never blended into a single number. Vendor self-reported figures (ABBYY's 99.8%) are labeled as vendor-claimed and used only as upper-bound references, per the D-grade citation rule requiring corroboration — in this case by IntuitionLabs' independent measurement that all engines exceed 99.2% on easy typed text.

Source List

  1. Holley, R. — "How Good Can It Get?" D-Lib Magazine 15(3/4) (2009). National Library of Australia newspaper digitization program; raw character-confidence measurements across 30+ historic newspaper titles, 1803–1954. Provides the 71–98.02% degraded-scan range, the 98–99%/90–98%/<90% good/average/poor scale, and the 1–2% CER "good" benchmark.
  2. Hebert, J. et al. — "OCR with Tesseract, Amazon Textract, and Google Document AI: a benchmarking experiment," Journal of Computational Social Science (2021). Peer-reviewed; n=322 English book scans + n=100 Arabic article scans × 43 noise types = 18,568 documents, 51,304 requests. Provides the clean-to-noisy word-accuracy drop (~96%→82.5%→58.4%), English-vs-Arabic gap, and server-vs-open-source comparisons.
  3. Rice, S.V., Jenkins, F.R. & Nartker, T.A. — "The Fourth Annual Test of OCR Accuracy," UNLV ISRI (1995). Annual independent evaluation of commercial OCR on 460-page DOE sample + 200-page magazine sample. Provides the 1992–1995 character-accuracy progression (~94%→98%+) that anchors the clean-text ceiling.
  4. NIST IR 6101 — "Impact of image quality on machine print optical character recognition" (1998). US National Institute of Standards and Technology evaluation of three commercial OCR products on Federal Register pages. Provides the "1% error rate requires fixed-format clean originals" constraint.
  5. European Data Protection Board — "AI Possible Risks & Mitigations: OCR" (2024). EU regulatory body technical guidance. Provides the 95–99% common printed-document range and the 98–99% production guarantee level.
  6. US Government Publishing Office — "Optimizing OCR Accuracy on Older Documents" white paper. Official testing of scan modes and enhancements on aged/discolored documents against the GPO 99% accuracy requirement. Provides the RGB 99% vs bitonal ~77% finding and the 300→200 DPI downsampling result.
  7. "Optical Character Recognition Accuracy on Degraded Documents," IJSAT (2026). Peer-reviewed; EasyOCR on four physical conditions at 300 DPI. Provides the 94.7% normal / 86.9% crumpled / 80.9% dirty / 72.7% wet accuracy ladder.
  8. AIMultiple — "OCR Benchmark: Text Extraction / Capture Accuracy" (2026). Vendor-neutral analyst benchmark using Sentence-BERT cosine similarity; 12+ solutions. Provides the 46–95% handwriting band and per-solution leaders (GPT-5 95%, olmOCR-2-7B 94%, Gemini 2.5 Pro 93%).
  9. AIMultiple — "Invoice OCR Benchmark: Extraction Accuracy of LLMs vs OCRs" (2026). Vendor-neutral benchmark, 400+ key-value pairs from 20 publicly available invoices. Provides the low-quality-vs-high-quality accuracy decline finding.
  10. Gao, L. et al. — "ICDAR 2019 Competition on Table Detection and Recognition (cTDaR)," ICDAR (2019). Academic competition, 11 teams, modern + archival datasets with hand-drawn and handwritten tables. Provides the archival track wF1 0.49 (later models 0.53) table-detection floor.
  11. OpenCodePapers — "Table Recognition on PubTabNet" leaderboard (2024). Technical benchmark aggregator tracking published TEDS results on the PubTabNet corpus. Provides the TEDS ceiling: MuTabNet 0.9687 (2024), TableMaster 0.9676, SLANet 0.963.
  12. "Tabular context-aware OCR and tabular data reconstruction for historical records" (2025). Peer-reviewed; TrOCR-ctx across PubTabNet, SROIE, CORD, historical corpora. Provides TEDS 0.909 (PubTabNet), 0.967 (SROIE), 0.986 (CORD) table-structure scores.
  13. "Enhancing handwritten text recognition accuracy with gated mechanisms" (2024). Peer-reviewed; GatedLexiconNet on IAM, RIMES, READ-2016. Provides IAM WER 5.73%/CER 2.27% and RIMES CER 0.9% state-of-the-art handwriting figures.
  14. "Benchmarking Large Language Models for Handwritten Text Recognition" (2025, Journal of Documentation). Peer-reviewed; proprietary and open LLMs vs Transkribus across 7 corpora. Provides GPT-4o-mini IAM 1.71% CER/3.34% WER, GPT-4o RIMES 1.69% CER/3.66% WER, and the modern-CER<5% quality framing.
  15. BusinessWareTech — "AWS Textract vs Google, Azure, and GPT-4o: Invoice Extraction Benchmark" (2025). Independent benchmark on real invoices; C-grade (single test set, transparent method). Provides field accuracy 78–98% by engine and line-item detection 40–82% by engine.
  16. IntuitionLabs — "Pharma Document AI & OCR Accuracy: A Benchmark Report" (2025-04). Industry research summarizing a mixed-document benchmark; C-grade. Provides the all-engines >99.2% on easy typed text finding and Google Cloud Vision ~98% overall.
  17. ABBYY FineReader (vendor claim). D-grade self-reported figure, used only as the clean-text upper bound and labeled as such. "Up to 99.8% accuracy on clean scans"; also documents the fax/low-resolution improvement claim.
  18. LlamaIndex — "What is OCR Accuracy Rate?" glossary (2026). Reference glossary with factor-impact table. Provides the DPI ladder (72 vs 300 vs 600 DPI) and 300-DPI minimum guidance used in the DPI table.
  19. University of Pittsburgh — "Best Practices: OCR" library guide. University digitization guidance. Provides the 300 DPI standard and 400–600 DPI for <10pt fonts recommendation.
  20. Cassidy, S. — "The Impact of OCR Quality on the Use of Digitized Historical Texts" (Macquarie University). Peer-reviewed study of post-OCR correction on historical newspapers. Provides the 81.7%→97.6% word-accuracy correction effect.

Limitations

  • Geographic coverage: Sources cluster in North America, the EU, Australia, and (for handwriting) European corpora. No reliable primary data was found for Latin America or Africa; CJK/Devanagari comparisons are absent from the peer-reviewed literature we located.
  • Methodology constraints: Cloud vendors (Google, Azure, Amazon) publish no methodology-transparent accuracy benchmarks by document type — their "99%+" figures are marketing, not measurements, and are excluded here except where an independent study measured them. Several key figures rely on single-study data points (BusinessWareTech field accuracy; IJSAT degradation) and are labeled with their grade. ABBYY's 99.8% is a vendor claim used only as a corroborated upper bound.
  • Temporal gaps: The UNLV ISRI (1995) and NIST (1998) data are 27+ years old but remain the only public longitudinal records of the clean-text accuracy ceiling; they are labeled by year and used for historical context. The Holley (2009) degraded-scan floor is the most recent available published raw figure for historic newspapers — modern engines score better on such documents (Hebert et al. 2021), which is why the table shows both.
  • What we could not find: (1) any reliable public benchmark for checkbox/mark-sense recognition in modern AI OCR — the literature is legacy OMR; (2) signature recognition accuracy — vendors report detection, not recognition accuracy; (3) a peer-reviewed controlled study isolating DPI while holding other factors constant; (4) standardized benchmarks for screen-capture (screenshot) OCR; (5) mixed-layout documents (text + tables + images on one page) beyond the line-item figures, which are the only published proxy. These gaps are stated rather than filled with unverifiable vendor numbers.

Related references: Manual Data Entry Error Rates: 1-4% by Type, Industry & Role · What is Optical Character Recognition (OCR)?

Related reading: How to Improve OCR Accuracy: 10 Practical Tips · How Accurate Is AI Document Extraction Really? · ABBYY FineReader vs Modern AI OCR

📮 contact email: [email protected]