Why CER Misleads for Document-Parsing VLMsQuantified with First-Party Benchmark Data (2026)

Last reviewed: 2026-08-18 · Run tier: official · First-party benchmark · 8 engines × 2 receipt datasets

What this page covers: A first-party, reproducible quantification of when and how far Character Error Rate (CER) misleads when comparing traditional OCR engines against document-parsing vision-language models (VLMs) — the two-layer mechanism (VLM output normalization, then ground-truth structure inflation), the measured CER-vs-field-F1 ranking inversions, and the metrics that stay fair across the family boundary. Every figure traces to a published CSV row in the public OCR benchmark repository (ImageToTableai/benchmark-ocr); the ~76%/~18%/~10% CER-decomposition figures are benchmark protocol analysis estimates, not CSV columns, and are labeled as such wherever they appear.
What this page does NOT cover: The definitional difference between character-level and field-level accuracy — that lives on the companion field vs character accuracy definition page; this page is the data layer that quantifies the gap. No invoices, forms, or long documents are measured — receipts only (SROIE 2019 English, CORD v2 Indonesian), one GPU tier (RTX 4090), August 2026 model versions.

Range statement: Mechanism illustrated on receipts (SROIE 2019 English, 361 test samples; CORD v2 Indonesian, 100 test samples) using the first-party benchmark data. The ~76%/~18%/~10% CER decomposition is protocol-note-derived estimate, not a CSV column. CORD is never merged into the SROIE ranking — different language, different ground-truth structure.

Character Error Rate (CER) is an exact-character matching algorithm, not a quality meter: it charges one error for every character that differs from ground truth. In this benchmark, that scoring convention alone — before any genuine misreading — moves engines between the top and bottom of the ranking: Unlimited-OCR scores the worst CER (0.6552) and the best regex field F1 (0.3376) on the same 361 SROIE receipts, while docTR scores the 2nd-best CER (0.1971) and the worst field F1 (0.0766).

The three numbers writers most often need: ~18% of a VLM’s SROIE CER budget is case substitutions alone (protocol-derived estimate) — folding TAN CHAY YEE to tan chay yee is scored as errors even though the field value is correct; 1.0805, PaddleOCR-VL’s CORD CER — the worst of all 8 engines, an artifact of CORD’s structure-laden ground truth, while its CORD field F1 of 0.3412 is the benchmark’s best; and the 0.6552 / 0.3376 CER-worst / field-best inversion that makes ranking by CER alone pick the wrong winner.

~18%
Share of a VLM’s SROIE CER budget attributed to case substitutions in the benchmark’s error-decomposition analysis — characters are correct, only the case is folded (protocol-derived estimate from the benchmark’s analysis notes, not a CSV column)
1.0805
PaddleOCR-VL’s CORD CER — the worst of all 8 engines on a metric that embeds annotation structure in the ground truth; the same engine’s CORD regex field F1 (0.3412) is the benchmark’s best (summary_metrics.csv, cer / field_f1_regex, paddleocr_vl_vllm/cord_v2 row)
0.6552 ↔ 0.3376
Unlimited-OCR’s CER (worst of 8) versus its regex field F1 (best of 8) on the same SROIE receipts — the sharpest ranking inversion in the benchmark (summary_metrics.csv, cer / field_f1_regex, unlimited_ocr/sroie_2019 row)

CER Is an Exact-Match Scoring Algorithm, Not a Quality Meter

CER is the Levenshtein edit distance between recognized and ground-truth text — the minimum number of character insertions, deletions, and substitutions needed to turn one into the other, divided by the ground-truth length (the formal definition is published by the OCR-D evaluation specification). It counts differences, and it cannot tell a misreading apart from a legitimate reformatting. When an engine’s output convention differs from the ground truth — case, separators, line order — CER charges errors for behavior that is not misreading at all.

Traditional OCR engines (Tesseract, PaddleOCR, EasyOCR, docTR) emit raw character streams that preserve original casing and line layout, so their output convention sits close to the ground truth and CER measures something near a genuine reading error. Document-parsing VLMs (Surya2, Unlimited-OCR, PaddleOCR-VL) emit understood text: they apply case folding (TAN CHAY YEE becomes tan chay yee), label/value merging (INVOICE NO\n: PEGIV becomes Invoice No : PEGIV), and line reordering — the output conventions of a reader, not a scanner. Every folded case and merged line is an edit-distance penalty, even when the underlying field value is correct.

The benchmark’s protocol analysis of the released SROIE predictions decomposes a VLM’s raw CER budget into roughly ~76% characters identical to ground truth, ~18% case substitutions, and ~10% line-merges / dropped separators — with the actual field values (company, total, date, address) correct. This decomposition is a protocol-derived estimate from the benchmark’s analysis notes, not a CSV column; treat the split as directional, not precise. The direction is what matters: case folding alone can explain a large chunk of a VLM’s “bad” CER on clean English receipts.

Two CSV rows make the mechanism visible without any decomposition. Unlimited-OCR scores CER 0.6552 but WER 0.4779 on SROIE — its characters look mangled while its words survive, because case folding replaces characters without breaking words (summary_metrics.csv, cer / wer, unlimited_ocr/sroie_2019 row). And the VLM that normalizes least, Surya2, scores CER 0.1915 — the best raw CER in the entire 8-engine run, indistinguishable from a strong traditional engine (summary_metrics.csv, cer, surya2/sroie_2019 row). VLM output convention, not VLM reading ability, is what the CER column is mostly measuring.

The Proof Is in the Inversions: CER and Field F1 Disagree

The benchmark’s CER ranking and its field-extraction ranking disagree materially. Rank the eight engines by SROIE CER and the winner is Surya2 (0.1915); rank them by regex field F1 — the metric that approximates what a production pipeline consumes — and the winner is Unlimited-OCR (0.3376), the engine CER ranks dead last. One of the two columns is not measuring what downstream systems actually consume.

Field F1 is the harmonic mean of precision and recall over extracted field values (company, date, address, total on SROIE): a field is a binary unit — it matches or it fails. The regex column applies the same fixed pattern set to every engine’s text, so the only variable is the engine’s output. The table below pairs each engine’s CER with its regex field F1 on the same 361 receipts, with each engine’s rank under both metrics.

SROIE 2019 CER vs regex field F1 by engine: CER (lower is better) — Surya2 0.1915, docTR 0.1971, PaddleOCR 0.2045, EasyOCR 0.2833, Tesseract 0.3347 (CPU), PaddleOCR-VL 0.3370, Docling 0.5909, Unlimited-OCR 0.6552. Regex field F1 (higher is better) — Unlimited-OCR 0.3376, PaddleOCR-VL 0.3368, PaddleOCR 0.3254, Surya2 0.3183, Tesseract 0.2335, Docling 0.2237, EasyOCR 0.1477, docTR 0.0766.

Source: summary_metrics.csv — cer and field_f1_regex columns, sroie_2019 rows (361 samples per engine, error_rate 0.0 for all 8). The two series are not comparable to each other as magnitudes (different units), but their rankings disagree, which is the point.

ModelTypeCERWERRegex field F1CER rankField-F1 rankSource
Surya2Document-parsing VLM0.19150.27350.318314summary_metrics.csv · surya2/sroie_2019 row
docTRTraditional OCR (GPU)0.19710.31990.076628summary_metrics.csv · doctr/sroie_2019 row
PaddleOCRTraditional OCR (GPU)0.20450.32560.325433summary_metrics.csv · paddleocr/sroie_2019 row
EasyOCRTraditional OCR (GPU)0.28330.61580.147747summary_metrics.csv · easyocr/sroie_2019 row
TesseractTraditional OCR (CPU)0.33470.55910.233555summary_metrics.csv · tesseract/sroie_2019 row
PaddleOCR-VLDocument-parsing VLM0.33700.64620.336862summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row
DoclingPipeline parser0.59090.75960.223776summary_metrics.csv · docling/sroie_2019 row
Unlimited-OCRDocument-parsing VLM0.65520.47790.337681summary_metrics.csv · unlimited_ocr/sroie_2019 row

Table: summary_metrics.csv — cer / wer / field_f1_regex columns, sroie_2019 rows. Ranks computed within the 8 rows of this table (1 = best under that metric: lowest CER, highest field F1). These are postprocessed_sroie_receipt_regex_* metrics: fixed patterns applied to each engine’s OCR text, not native structured output. Tesseract ran CPU-only (compute_type=cpu).

Read the two inversion pairs explicitly. Unlimited-OCR: worst CER (0.6552), best field F1 (0.3376) — the engine the raw-OCR metric ranks last is the engine the field lens ranks first. docTR: 2nd-best CER (0.1971), worst field F1 (0.0766) — text-perfect, field-fail. PaddleOCR-VL inverts at the margin too (CER rank 6, field rank 2), while the two aligned rows (PaddleOCR 3rd/3rd, Tesseract 5th/5th) are both traditional engines whose raw-text output convention is exactly what CER was designed to score. Rank engines by CER alone and you get a different winner than ranking by what production consumes — the inversion is the demonstration, not an anomaly.

The Field Metric Is the Reliable Cross-Family Lens

Run each engine’s raw text through the same LLM field-extraction postprocessor (deepseek-v4-flash, temperature 0) and the six engines with usable OCR text converge to 0.57–0.62 field F1 — the VLM-vs-traditional family boundary that the CER ranking makes look enormous (0.19 to 0.66) nearly disappears. Two engines fall below the band: EasyOCR at 0.3717 and Tesseract at 0.4389. The field metric separates engines by what actually matters — downstream field recovery — and it does so consistently across the family boundary.

This is the same engine set as the CER table above, re-scored on the field dimension only: the CER inversion between Unlimited-OCR and docTR is gone, because field value recovery is what the pipeline consumes. The mechanism is that the postprocessor absorbs the output-convention differences CER punished — it reads the case-folded, label-merged text and extracts the values. LLM field F1 here is the llm_field_value_f1 column of the benchmark’s method comparison, with the LLM model recorded per row.

ModelTypeCER (context)LLM field F1In convergence band (0.57–0.62)Source
docTRTraditional OCR (GPU)0.19710.6171Yes — highestfield_method_comparison.csv · llm_field_value_f1, doctr/sroie_2019 row
Surya2Document-parsing VLM0.19150.6139Yesfield_method_comparison.csv · llm_field_value_f1, surya2/sroie_2019 row
Unlimited-OCRDocument-parsing VLM0.65520.6054Yesfield_method_comparison.csv · llm_field_value_f1, unlimited_ocr/sroie_2019 row
PaddleOCR-VLDocument-parsing VLM0.33700.5921Yesfield_method_comparison.csv · llm_field_value_f1, paddleocr_vl_vllm/sroie_2019 row
PaddleOCRTraditional OCR (GPU)0.20450.5810Yesfield_method_comparison.csv · llm_field_value_f1, paddleocr/sroie_2019 row
DoclingPipeline parser0.59090.5685Yes — band edgefield_method_comparison.csv · llm_field_value_f1, docling/sroie_2019 row
TesseractTraditional OCR (CPU)0.33470.4389No — belowfield_method_comparison.csv · llm_field_value_f1, tesseract/sroie_2019 row
EasyOCRTraditional OCR (GPU)0.28330.3717No — belowfield_method_comparison.csv · llm_field_value_f1, easyocr/sroie_2019 row

Table: field_method_comparison.csv — llm_field_value_f1 column, sroie_2019 rows, llm_model=deepseek-v4-flash. CER column repeated from summary_metrics.csv for cross-reference only. The “convergence band” is the observed 0.5685–0.6171 range of the six engines with usable text; the two below-band rows are stated as observations, not a classification of the engines.

CORD Adds a Second Inflation Layer: Ground-Truth Structure

CORD’s published ground-truth text (gt_text) embeds annotation structure — field labels, menu entries, and coordinates — rather than pure visible text. CER computes edit distance against that structurally augmented string, so every engine’s CORD CER is systematically inflated before any reading error is even considered. The result: all eight engines cluster at CORD CER 0.90–1.08 — a uselessly compressed range that says almost nothing about text quality.

The extreme artifact is PaddleOCR-VL’s CORD CER of 1.0805, the worst of all 8 engines — not a reading of its text quality but the structural inflation working hardest against the cleanest, shortest output, which has the largest relative edit distance to CORD’s structure-laden ground truth. On the field lens, the same engine posts the benchmark’s best CORD regex field F1 (0.3412). CORD CER is structurally unusable; field metrics are the only fair CORD lens, and they are kept strictly separate from any SROIE ranking in this benchmark (different language — Indonesian — and different ground-truth structure).

CORD v2 CER vs regex field F1 by engine: CER (lower is better) — Surya2 0.8959, PaddleOCR 0.9083, docTR 0.9101, EasyOCR 0.9185, Docling 0.9219, Unlimited-OCR 0.9224, Tesseract 0.9523 (CPU), PaddleOCR-VL 1.0805. Regex field F1 (higher is better) — PaddleOCR-VL 0.3412, Surya2 0.2458, Unlimited-OCR 0.1079, Tesseract 0.0752, Docling 0.0612, PaddleOCR 0.0154, EasyOCR 0.0067, docTR 0.0.

Source: summary_metrics.csv — cer and field_f1_regex columns, cord_v2 rows (100 samples per engine). CORD CER is not comparable across families or even across engines — the ground truth embeds annotation structure, so CER measures distance to a structurally augmented string, not to visible text.

ModelTypeCORD CERRegex field F1Source
PaddleOCR-VLDocument-parsing VLM1.08050.3412summary_metrics.csv · paddleocr_vl_vllm/cord_v2 row
Surya2Document-parsing VLM0.89590.2458summary_metrics.csv · surya2/cord_v2 row
Unlimited-OCRDocument-parsing VLM0.92240.1079summary_metrics.csv · unlimited_ocr/cord_v2 row
TesseractTraditional OCR (CPU)0.95230.0752summary_metrics.csv · tesseract/cord_v2 row
DoclingPipeline parser0.92190.0612summary_metrics.csv · docling/cord_v2 row
PaddleOCRTraditional OCR (GPU)0.90830.0154summary_metrics.csv · paddleocr/cord_v2 row
EasyOCRTraditional OCR (GPU)0.91850.0067summary_metrics.csv · easyocr/cord_v2 row
docTRTraditional OCR (GPU)0.91010.0000summary_metrics.csv · doctr/cord_v2 row

Table: summary_metrics.csv — cer / field_f1_regex columns, cord_v2 rows. CORD is not merged into any SROIE ranking (protocol rule): the ground-truth structure inflates CER for every engine, and the regex patterns were written for English formats. Field metrics are the only fair CORD lens. Sorted by field F1, not CER — the point is that the CER column carries no sorting signal.

The field lens flips the CORD table completely. The engine with the benchmark’s worst CORD CER (PaddleOCR-VL, 1.0805) has its best CORD field F1 (0.3412); docTR’s CORD field F1 is literally 0.0000 — nothing extracted — while its CORD CER (0.9101) sits mid-cluster and says nothing about it. On CORD, publishing CER without the field numbers is not just uninformative; it actively ranks engines in the wrong order.

When CER Is Actually the Right Metric

CER still measures something real — exact-character reproduction — and it is the right metric whenever the downstream consumer needs verbatim text, not fields. The conclusion of this page is “CER misleads for cross-family comparison of VLM-style outputs,” not “CER is always wrong.”

  • Within the raw-OCR family, CER retains its value. Traditional engines emit raw text with preserved casing and layout, so CER measures genuine reading quality: the traditional-engine CER column in this benchmark (0.1971–0.3347 on SROIE) tracks their field performance far more closely than it does for the VLMs (rank correlation with regex field F1: PaddleOCR and Tesseract are exactly aligned at 3rd/3rd and 5th/5th).
  • Verbatim-text use cases. Exact-match downstream consumers — full-text search that must find a literal string, audit trails that replay a document’s characters, or pass/fail checks against a required character sequence — consume characters, not fields. For those, exact-character reproduction (CER) is the faithful measure and field F1 is the wrong lens.
  • Publication rule of thumb. When you publish a CER number, state both the engine’s output convention (does it fold case? merge labels? reorder lines?) and the ground-truth construction (pure visible text, or annotation-augmented like CORD’s gt_text). If either is unknown, the number is not portable across engines or datasets.
  • The boundary rule from this benchmark. CER is fair within the raw-OCR family and unfair across the traditional-OCR / document-parsing-VLM boundary; field-level F1 (regex or LLM postprocessed) is fair across both. Use CER for same-family raw text quality, field F1 for cross-family comparison — and always report the decomposition caveat for VLM rows.

Frequently Asked Questions

Why is CER misleading when comparing OCR engines and VLMs?

Because CER is an exact-character edit-distance score, and document-parsing VLMs deliberately do not reproduce characters exactly — they fold case, merge label/value lines, and reorder text. Every one of those normalizations is scored as an error even when the field value is correct. In this benchmark, Unlimited-OCR scores the worst SROIE CER (0.6552) while having the best regex field F1 (0.3376) on the same 361 receipts (summary_metrics.csv, cer / field_f1_regex, unlimited_ocr/sroie_2019 row) — the CER ranking and the field ranking pick different winners.

Is CER still a useful OCR metric at all?

Yes — within the raw-OCR family and for verbatim-text consumers. Traditional engines (Tesseract, PaddleOCR, EasyOCR, docTR) output raw character streams that CER measures fairly: in this benchmark their SROIE CER spans 0.1971–0.3347 and tracks field performance (PaddleOCR 3rd/3rd, Tesseract 5th/5th under CER and field F1). CER is the wrong metric when the engine normalizes its output (VLM-style) or when the ground truth embeds structure rather than visible text (CORD).

Why do VLM OCR models have high character error rates?

Mostly output convention, not reading ability. The benchmark’s protocol analysis attributes roughly ~18% of a VLM’s SROIE CER budget to case substitutions and ~10% to line-merges/dropped separators (protocol-derived estimate from the benchmark’s analysis notes, not a CSV column) — while the field values themselves are correct. Unlimited-OCR’s CER 0.6552 vs WER 0.4779 shows the signature: characters get folded, words survive (summary_metrics.csv, cer / wer, unlimited_ocr/sroie_2019 row).

What is the best metric to compare OCR engines and document-parsing VLMs?

Field-level F1 — regex postprocessed for a raw-OCR pipeline, LLM postprocessed for either family. On SROIE, LLM field F1 converges six engines to 0.57–0.62 regardless of family (field_method_comparison.csv, llm_field_value_f1, sroie_2019 rows), and regex field F1 is the only lens that ranks the CORD table sensibly, where PaddleOCR-VL’s 0.3412 is the benchmark’s best (summary_metrics.csv, field_f1_regex, cord_v2 rows). Always state the postprocessor and the output convention alongside the number.

Why is PaddleOCR-VL’s CER so high on CORD but its field F1 the best?

Because CORD’s ground truth embeds annotation structure (field labels, coordinates, menu entries) rather than pure visible text, and PaddleOCR-VL’s output is the cleanest — shortest, most normalized — so its edit distance to that structurally augmented string is the largest (CER 1.0805, worst of 8). On the field lens, the same text extracts Indonesian receipt fields at 0.3412 field F1, the benchmark’s best (summary_metrics.csv, cer / field_f1_regex, paddleocr_vl_vllm/cord_v2 row). CORD CER is structurally unusable for every engine; field metrics are the only fair CORD comparison.

What is case folding and why does it inflate CER?

Case folding is normalizing all text to a single case — TAN CHAY YEE becomes tan chay yee. CER compares characters exactly, so every folded letter is a substitution against the ground truth’s uppercase, even though the value is identical. In this benchmark’s protocol analysis, case substitutions account for roughly ~18% of a VLM’s SROIE CER budget (protocol-derived estimate) — which is why a VLM can read a receipt perfectly and still report a CER that looks like a failing model.

Should I compare OCR engines by CER or by field-level F1?

By field-level F1 for any comparison that touches both traditional OCR and VLM engines — because your pipeline consumes fields, not character streams, and because CER is systematically inflated for normalized VLM output and structure-augmented ground truth. Use CER only within the raw-OCR family or when the downstream consumer needs verbatim text (search, audit, exact-passing). The two metrics disagree sharply in this benchmark: CER ranks docTR 2nd and Unlimited-OCR 8th; regex field F1 ranks docTR 8th and Unlimited-OCR 1st (summary_metrics.csv, sroie_2019 rows).

Where do these CER and field F1 numbers come from?

Every table number is a row of the first-party benchmark’s published results/summary_metrics.csv or results/field_method_comparison.csv hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run recording model versions and environment fingerprints. The ~76%/~18%/~10% CER decomposition is a protocol-derived estimate from the benchmark’s analysis notes, explicitly not a CSV column. Dataset definitions come from the SROIE and CORD papers cited below.

Methodology & Sources

Protocol

This page reports the metric-fairness dimension of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total, CC-BY-4.0) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total, CC-BY-4.0); training splits were never evaluated. Eight engines ran out-of-the-box, no fine-tuning, under the frozen protocol reports/receipt_v1_official_protocol.md (warm-then-scored, official tier only). All 16 scored runs completed with error_rate 0.0 (summary_metrics.csv error_rate column). Data was collected in August 2026.

Runtime Environment

  • Hardware: all GPU runs on a single NVIDIA RTX 4090 (24 GB); Tesseract ran CPU-only (compute_type=cpu) and is labeled as such on every table.
  • Engines and versions (per run manifests): Tesseract 5.3.4, PaddleOCR 3.7.0, EasyOCR 1.7.2, docTR v1.0.1, Docling 2.119.0, Surya2 0.22.1, Unlimited-OCR vLLM-served, PaddleOCR-VL 1.6.
  • Field postprocessors: regex field F1 = fixed sroie_receipt_regex/cord_receipt_regex patterns applied to each engine’s OCR text (postprocessed, not native extraction); LLM field F1 = deepseek-v4-flash (temperature 0) reading the same text (llm_model column of field_method_comparison.csv).

Metric Definitions

  • Character Error Rate (CER): Levenshtein edit distance between recognized and ground-truth text — minimum insertions/deletions/substitutions divided by ground-truth length (OCR-D evaluation specification). Lower is better. Measures exact-character reproduction only.
  • Word Error Rate (WER): the same edit-distance logic at word granularity — a word survives a single case substitution, so CER/WER divergence reveals normalization that changes characters without breaking words.
  • Field-value F1: harmonic mean of precision and recall over extracted field values; a field matches only if it exactly equals ground truth. The regex and LLM columns are two postprocessors over the same OCR text, never mixed.
  • Output convention: how an engine formats its text (case, separators, line order). Traditional engines preserve it; document-parsing VLMs normalize it — the axis this page measures.
  • Ground-truth structure: whether the reference text is pure visible text (SROIE) or augmented with annotation structure (CORD gt_text), which inflates CER for every engine.

Source List

  1. summary_metrics.csv (GitHub raw). 16 rows = 8 engines × 2 receipt datasets. Columns include cer, wer, field_f1_regex, field_acc_regex, compute_type, error_rate. Every CER/WER and regex field F1 figure on this page traces to a row here, cited at file/model/dataset/metric level.
  2. field_method_comparison.csv (GitHub raw). Regex vs LLM field extraction side by side (llm_field_value_f1, llm_model=deepseek-v4-flash). Source for the LLM field F1 table.
  3. ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, the frozen protocol (reports/receipt_v1_official_protocol.md), and fixed sample lists for reproduction.
  4. results/manifests/ (GitHub). One redacted manifest.json per published run with model versions, environment fingerprints, and artifact hashes.
  5. receipt_v1_official_protocol.md. The frozen run contract: fixed test splits, official tier, warm-then-scored measurement, CORD-vs-SROIE separation rules. Source for the protocol-level rules quoted on this page.
  6. Benchmark protocol analysis notes (internal run documentation, WRITING_BRIEF §8.2/§8.3). The VLM normalization mechanism (case folding, label/value merging, line reordering), the CORD ground-truth structure mechanism, and the SROIE CER decomposition (~76% identical / ~18% case / ~10% line-merge). Cited here with the explicit label “protocol-derived estimate — not a CSV column”; the split is directional, not precise.
  7. OCR-D Project — Quality Assurance Specification. Formal CER definition (insertions + deletions + substitutions) / total characters and the WER analogue. Formal definitional anchor.
  8. Huang et al., “ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction” (2019). SROIE 2019 dataset definition, task structure, license (CC-BY-4.0).
  9. Park et al., “CORD: A Consolidated Receipt Dataset for Post-OCR Parsing” (2020). CORD dataset definition, nested field schema, license (CC-BY-4.0).

Limitations

  • Decomposition figures are protocol-derived estimates, not measurements: the ~76%/~18%/~10% SROIE CER decomposition comes from the benchmark’s protocol analysis notes (WRITING_BRIEF §8.2), not from a CSV column. The direction (normalization, not misreading, dominates VLM CER) is robust; the exact percentages should be treated as directional estimates, not point measurements.
  • Receipts only: SROIE (English) and CORD (Indonesian) receipts. Behavior on invoices, forms, tables, or long documents is unmeasured; the mechanism generalizes, the numbers do not.
  • Single LLM postprocessor: all LLM field F1 figures use deepseek-v4-flash at temperature 0. A different LLM shifts the absolute F1 values and possibly the convergence band; the cross-family ordering within the band is the stable signal.
  • CER still valid for verbatim-text cases: the conclusion is scoped to cross-family comparison of VLM-style outputs. For raw-OCR same-family comparison and exact-character downstream consumers (search, audit), CER remains a legitimate metric — this page is not an argument against CER in those contexts.
  • CORD not merged into SROIE rankings: different language, different ground-truth structure. CORD is compared on field metrics only, per the benchmark protocol.
  • Version and hardware pinning: numbers hold for the August 2026 model versions and one GPU tier (RTX 4090) listed above. Newer releases shift CER and field F1; single-digit-percent differences should be treated as noise.
  • Sample size: 361 + 100 samples; bootstrap confidence intervals are reported in the benchmark’s evaluation output but are not reproduced row-by-row on this page.

Related references: Field-Level vs Character-Level Accuracy (definition) · Traditional OCR vs Document Parsing VLMs · Regex vs LLM Field Extraction · PaddleOCR vs EasyOCR on Receipts · Surya2 vs Unlimited-OCR vs PaddleOCR-VL · OCR Latency Benchmark · OCR Cost per 1,000 Pages

Related reading: why AI OCR accuracy diverges from traditional OCR · AI extraction from images versus traditional OCR pipelines

📮 contact email: [email protected]