Field-Level vs Character-Level Accuracy: What's the Difference?
Last reviewed: 2026-08-13 · Applies to: document data extraction / OCR evaluation / AP and data-entry workflows
Also known as: Field-level accuracy is sometimes called "field accuracy," "field extraction accuracy," or "field-level extraction accuracy." Character-level accuracy is usually expressed as its inverse, Character Error Rate (CER).
Field-level accuracy measures whether each extracted field is completely correct; character-level accuracy measures whether individual characters are read correctly. One misread character corrupts an entire field, so the two metrics describe different things entirely.
How the Two Accuracy Metrics Work
Both metrics are ratios, but they count different things. Character-level accuracy asks "how many of the characters on this document were read correctly?" Field-level accuracy asks "how many of the data fields I actually need were extracted without a single error?"
Character-level accuracy = correctly recognized characters ÷ total characters. The standard formal definition comes from the OCR evaluation community: Character Error Rate (CER) is the Levenshtein edit distance between the recognized text and the ground truth — the minimum number of substitutions, deletions, and insertions needed to turn one into the other — divided by the total character count (OCR-D project specification). Character accuracy is simply 100% − CER. NIST's own evaluation framework (TRAIT) defines Character Accuracy the same way, as 100 × (1 − normalized edit distance) over all annotations (NISTIR 8199, 2017).
Field-level accuracy = correctly extracted fields ÷ total fields. A field — invoice total, date, vendor name, account number — is scored as a binary unit: either it exactly matches the ground truth or it fails. The formal version used in academic benchmarks is field-level F1, the harmonic mean of precision and recall: F1 = 2 × (Precision × Recall) / (Precision + Recall), where precision is the share of extracted fields that were correct and recall is the share of expected fields that were found (LlamaIndex glossary, worked 10-field example). A single wrong digit in an invoice number makes that field wrong — partial credit does not apply.
This is where the two metrics diverge. A field with N characters at C% character accuracy has roughly C^N probability of being fully correct — the compounding effect that makes character accuracy useless for predicting field accuracy. For a 9-digit invoice number at 99% character accuracy: 0.99^9 ≈ 91.3% field accuracy for that field. The same number at 99.9% character accuracy rises to 99.1%. A 20-character field at 95% character accuracy drops to just 35.9% field accuracy. This is an approximation that assumes errors are independent and evenly distributed — real systems concentrate errors in the exact fields that matter (numbers, codes, identifiers), so actual field accuracy is usually worse than the model predicts. NIST's own OCR work observed this directly: recognizing page numbers produced a 12% field error rate even as overall character recognition looked respectable (NISTIR 6101).
The gap is not theoretical. On the ICDAR 2019 SROIE benchmark — 1,000 scanned receipts where teams were scored on both word-level OCR (Task 2) and field-level key information extraction (Task 3) on the same images — 7 of 24 submitted methods exceeded 90% word-level Hmean, but only 1 of 18 exceeded 90% field-level F1 (90.49%), and more than half of the field-level submissions scored below 80% (Huang et al., ICDAR 2019). Even the best OCR method in Task 2 could not meet the strict 99% accuracy that receipt applications demand — and field-level extraction was harder still. Modern field-level F1 on the CORD receipt benchmark sits between 78% and 97% depending on the model (LayoutLM 78.4, LayoutLMv2 78.9, Donut 84.1, LayoutLMv3 about 96.6), per the original model papers (Kim et al., Donut, 2022).
Why the Distinction Matters
Vendors quote character accuracy because it is the highest, easiest number to produce — while business decisions depend on field accuracy, which is always lower.
The marketing-to-reality gap is systematic. Character-level accuracy claims of 98–99.5% are routine for clean printed documents — and even the best OCR submissions to the ICDAR 2019 SROIE benchmark fell short of the 99% accuracy that receipt applications demand (Huang et al., ICDAR 2019). But field-level accuracy, measured on the same images, is far lower: only 1 of 18 field-level methods on SROIE exceeded 90% F1, while 7 of 24 word-level OCR methods cleared 90% — the same documents, two very different scores (Huang et al., ICDAR 2019). The C^N model above explains the gap: at 95% character accuracy, a 20-character field is fully correct just 35.9% of the time. Both numbers can be true on the same receipt — high character accuracy and far lower field accuracy — because they answer different questions.
The reason field accuracy is the number that matters is that your downstream systems consume fields, not characters. A wrong digit in an invoice total corrupts a payment; a wrong account number corrupts a posting — regardless of how many other characters were read correctly. NIST's own OCR technology demonstrated the point in its document-conversion research: recognizing Federal Register page numbers produced a 12% field error rate even as overall character recognition looked respectable (NISTIR 6101).
The operational consequence is review-queue sizing. Every field-level error is a document that must be caught and fixed. Modern models reach field-level F1 in the mid-90s on benchmark receipts — LayoutLMv3 scores about 96.6% on CORD, against 78–84% for earlier models (Kim et al., Donut, 2022) — but real-world accuracy depends on document type, image quality, and field design, and every document below 100% field accuracy routes someone into a review queue. A team evaluating a tool on character accuracy alone will consistently overestimate how many documents can flow through untouched — which is exactly why procurement and AP teams should request field-level accuracy on their own document types before committing to any pipeline.
Common Misconceptions
- Misconception: "99% accuracy means 99% of my data is correct."
- Reality: It depends entirely on what was measured. 99% character accuracy does not mean 99% field accuracy — a single misread character corrupts the whole field. At 99% character accuracy, a 9-digit invoice number has only a 91.3% chance of being fully correct (0.99^9), and a 20-character field is fully correct just 81.8% of the time. The same document can score high at the character level while a meaningful share of its fields are wrong — the SROIE benchmark documented 7 of 24 word-level methods above 90% against only 1 of 18 field-level methods (Huang et al., ICDAR 2019).
- Misconception: "Character accuracy and field accuracy are convertible — you can just subtract a fixed percentage."
- Reality: The conversion depends on field length and error distribution, so there is no fixed mapping. The C^N model shows the penalty grows with field length: a 9-character field at 99% character accuracy keeps 91.3% field accuracy, but a 20-character field at the same character accuracy drops to 81.8%, and at 95% character accuracy a 20-character field falls to just 35.9%. Different fields on the same document have different lengths and different error rates — no single conversion factor exists.
- Misconception: "A higher character accuracy always means a better tool."
- Reality: Field-level performance depends on more than raw character reading — context, structure understanding, and validation can lift field accuracy above what character accuracy predicts, or sink it below. A system can read nearly every character correctly and still assign a value to the wrong field — a structural error that character accuracy cannot even measure, since it never asks whether the right value landed in the right place. Field-level F1 catches this: it penalizes both wrong values and values extracted but mislabeled, which is why benchmarks score F1 rather than raw character accuracy (LlamaIndex glossary). Evaluate field accuracy directly, on your own documents.
Frequently Asked Questions
What is the difference between field-level and character-level accuracy?
Character-level accuracy measures how often individual characters are read correctly (as a ratio of correct characters to total characters); field-level accuracy measures how often complete data fields — invoice number, date, total, vendor — are extracted correctly as a unit. A field is wrong if even one of its characters is wrong, so field-level accuracy is always lower than character-level accuracy on the same document, and it is the metric that matters for real workflows.
Does 99% accuracy mean 99% of my data is correct?
Only if the 99% was measured at the field level. "99% accuracy" in vendor marketing is usually character-level accuracy, because that is the highest and easiest number to produce — and it is not convertible to field accuracy. At 99% character accuracy, a 9-digit invoice number is fully correct only about 91.3% of the time (0.99^9), and per-field error risk compounds across the fields of every document.
Which accuracy metric should I use when evaluating document extraction tools?
Ask for field-level accuracy, measured on your own document types. Character accuracy tells you how well the OCR engine reads text; it tells you nothing about whether the right value landed in the right field — the thing your downstream systems consume. When you see a headline accuracy claim, ask what was measured, on what documents, and on which fields. Field-level F1 (precision and recall together) is the standard formal metric because it also catches fields that were extracted but assigned the wrong label (LlamaIndex glossary).
How do you calculate field-level accuracy?
Field-level accuracy = correctly extracted fields ÷ total expected fields × 100%. A field matches only if it exactly equals the ground truth — one wrong character fails the field. The academic standard is field-level F1: F1 = 2 × (Precision × Recall) / (Precision + Recall), where precision is the share of extracted fields that were correct and recall is the share of expected fields that were found (LlamaIndex glossary). This is the same scoring the SROIE and CORD benchmarks use.
Why is field-level accuracy always lower than character-level accuracy?
Because a field fails if any one of its characters fails, and long fields fail quickly. The compounding model C^N means a 9-character field at 99% character accuracy keeps only 91.3% field accuracy, and the penalty grows with field length (0.99^20 ≈ 81.8%). Real systems also concentrate errors in numeric fields, where a single wrong digit is fatal — NIST observed this in its own OCR work, which produced a 12% field error rate on page numbers despite respectable character recognition (NISTIR 6101).
What are typical field-level accuracy benchmarks?
On academic benchmarks, field-level F1 ranges from 78% to 97% depending on the model: the ICDAR 2019 SROIE competition saw only 1 of 18 submissions exceed 90% field-level F1 (more than half scored below 80%), while the CORD benchmark's best models reach roughly 96–97% (Huang et al., 2019; Kim et al., 2022). Production numbers vary more than benchmarks because real documents are messier — treat any single figure as document- and field-dependent, and test on your own documents.
Sources
- Huang, Z., et al. — "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (IEEE ICDAR, 2019). 1,000-scanned-receipt benchmark with 24 OCR submissions and 18 field-level extraction submissions; 7 of 24 OCR methods exceeded 90% Hmean while only 1 of 18 field-level methods did (90.49%), over half below 80%. Primary source for the character-vs-field gap on identical documents.
- Kim, G., et al. — "OCR-free Document Understanding Transformer" (ECCV, 2022). CORD receipt benchmark field-level F1 across models: LayoutLM 78.4, LayoutLMv2 78.9, Donut 84.1, LayoutLMv3 (base) ~96.6. Source for the 78–97% field-level F1 range.
- NIST — "The Text Recognition Algorithm Independent Evaluation (TRAIT)" (NISTIR 8199, 2017). Official NIST definition of Character Accuracy = 100(1 − normalized edit distance)/#annotations. Source for the formal character-accuracy definition.
- NIST — "Impact of Image Quality on Machine Print OCR" (NISTIR 6101, 2001). NIST's own OCR produced a 12% field error rate on Federal Register page numbers despite respectable character-level recognition. Source for the field-error-rate example.
- OCR-D Project — Quality Assurance Specification. Formal CER = (insertions + deletions + substitutions) / total characters and the WER analogue. Source for the CER formula.
- LlamaIndex — "What is F1 Score for Document Extraction?". F1 = 2·(P·R)/(P+R) with a worked 10-field invoice example; explains field-level vs document-level scoring. Source for the F1 formula and worked example.
Related Terms
- What is Optical Character Recognition (OCR)?: The reading technology whose character-level accuracy is the source of the confusion — OCR turns pixels into characters; it does not by itself produce correct fields.
- character accuracy across document types: Character-level accuracy benchmarks broken down by document type — the layer that feeds (but never equals) field-level accuracy.
- What is Straight-Through Processing (STP) Rate?: The end-to-end metric that depends on field-level (and document-level) accuracy — a system needs near-100% field accuracy before documents can flow through with zero human touch.
Related reading: The 4-level accuracy guide: character, field, document, and STP · Why "99% accuracy" is nearly always character-level · How to measure field-level accuracy in practice