OCR Latency Benchmark
p50/p95, Throughput, and Tail Spikes (2026)
Last reviewed: 2026-08-18 · Run tier: official · First-party benchmark · 8 engines × 2 receipt datasets
What this page does NOT cover: Any document type other than receipts — no invoices, forms, contracts, or long documents. Cloud/API OCR services (AWS, Google, Azure), fine-tuned models, multi-GPU serving, batch-size sweeps, and CPU-only deployments beyond the Tesseract baseline are out of scope. The full 8-engine accuracy roundup lives on the OCR vs VLM accuracy comparison; the cost dimension of the same runs is on OCR Cost per 1,000 Pages.
Range statement: one GPU tier (RTX 4090) at one location/time — measure on your own hardware. Receipt datasets only (SROIE 2019, CORD v2). Steady-state p50/p95 exclude model loading; pages per minute are wall-clock including model init. p95 reflects first-page/prefill effects and batch-scheduling behavior under this run’s protocol (see Three Measures).
On the same GPU, the same receipts, and the same measurement protocol, median per-page OCR latency across 8 open-source engines spans 24.5× — from 108.7 ms (docTR) to 2,668.0 ms (Surya2) on SROIE 2019. Engine choice alone moves per-page latency by more than an order of magnitude on identical hardware — and the median hides the part that decides interactive experience: the tail.
The three numbers writers most often need: 108.7 ms p50 for the fastest engine measured (docTR, 449.3 pages/min) vs 2,668.0 ms p50 for the slowest (Surya2, 12.1 pages/min) on the same test split — and 26,976.9 ms (~27 s), Surya2’s p95 on CORD v2, the single worst tail event in the entire benchmark.
Three Measures, One Page: p50, p95, and pages/min
This page reports three measures of the same runs, and they answer different questions. p50 (the median) is steady-state per-page inference time in warm_then_scored mode — a fixed warm-up pass precedes the scored pass, and model loading is excluded — so it answers “how fast is this engine per page once it is already running?” p95 is the same measurement at the 95th percentile: 5% of pages took longer than this. Pages per minute is wall-clock throughput including model initialization — the number that governs batch jobs and billed GPU hours.
The three measures are not interchangeable and are not expected to agree. Steady-state p50 excludes model loading; wall-clock pages/min includes it; p95 captures first-page/prefill effects and batch-scheduling behavior that p50 ignores. Every chart and table on this page states which measure it shows — treat “108.7 ms” (p50, no load) and “449 pages/min” (wall-clock, with load) as two different facts about docTR, not a contradiction.
Three terms used throughout: tail latency is the behavior of the slowest few percent of pages (p95 and beyond) — for a user waiting on a single page, the tail, not the median, decides the experience. First-page/prefill effect is the one-time cost of priming a model or pipeline before steady inference, which shows up as slow early samples in a run. Wall-clock throughput counts every millisecond of the run, including initialization. On this benchmark’s protocol, p50/p95 are steady-state measurements and pages/min is wall-clock — the gap between them is init plus scheduling overhead.
SROIE 2019: Median Per-Page Latency Ranking
On 361 English receipts, the traditional two-stage OCR engines occupy the fast end and the document-parsing VLMs the slow end — but the spread inside each family is the surprise. docTR (108.7 ms) is 2.7× faster than the next traditional engine (PaddleOCR, 297.0 ms), and the two slowest engines are both VLMs (Unlimited-OCR 1,600.7 ms, Surya2 2,668.0 ms). Yet the fastest VLM, PaddleOCR-VL at 694.3 ms, still runs slower than every traditional engine.
Source: summary_metrics.csv — latency_p50_ms column, sroie_2019 rows. docTR 108.7, PaddleOCR 297.0, EasyOCR 413.6, Tesseract 670.9 (CPU-only), PaddleOCR-VL 694.3, Docling 732.0, Unlimited-OCR 1600.7, Surya2 2668.0. Steady-state latency, warm-then-scored measurement mode (excludes model loading).
| Rank | Model | Type | p50 (ms) | p95 (ms) | p95/p50 | Pages/min | Source |
|---|---|---|---|---|---|---|---|
| 1 | docTR | Traditional OCR (GPU) | 108.7 | 281.4 | 2.6× | 449.3 | summary_metrics.csv · doctr/sroie_2019 row |
| 2 | PaddleOCR | Traditional OCR (GPU) | 297.0 | 3,331.4 | 11.2× | 79.7 | summary_metrics.csv · paddleocr/sroie_2019 row |
| 3 | EasyOCR | Traditional OCR (GPU) | 413.6 | 960.4 | 2.3× | 124.5 | summary_metrics.csv · easyocr/sroie_2019 row |
| 4 | Tesseract | Traditional OCR (CPU) | 670.9 | 1,507.0 | 2.2× | 78.6 | summary_metrics.csv · tesseract/sroie_2019 row |
| 5 | PaddleOCR-VL | Document-parsing VLM | 694.3 | 1,154.3 | 1.7× | 68.2 | summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row |
| 6 | Docling | Pipeline parser | 732.0 | 3,239.8 | 4.4× | 56.7 | summary_metrics.csv · docling/sroie_2019 row |
| 7 | Unlimited-OCR | Document-parsing VLM | 1,600.7 | 2,521.9 | 1.6× | 34.4 | summary_metrics.csv · unlimited_ocr/sroie_2019 row |
| 8 | Surya2 | Document-parsing VLM | 2,668.0 | 5,872.2 | 2.2× | 12.1 | summary_metrics.csv · surya2/sroie_2019 row |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute, sroie_2019 rows (361 samples each, error_rate 0.0 for all 8). p95/p50 ratios computed by division of the CSV values (e.g., 3331.3505 / 296.9897 = 11.22). Tesseract ran CPU-only (compute_type=cpu); all others on GPU. Steady-state latency excludes model loading; pages/min is wall-clock including init.
The architecture pattern is clean at the family level — traditional engines occupy ranks 1–4, VLMs ranks 5, 7, 8, with Docling (a pipeline parser, not a pure OCR engine or a VLM) in the middle — but the gap within each family is wide. The VLM cluster spans 3.8× (694.3 to 2,668.0 ms) and the traditional cluster 6.2× (108.7 to 670.9 ms), which means “VLM” and “traditional” are not latency categories — individual engine design decides far more than family membership.
p50 Hides the Tail: p95/p50 Ratios Across Engines
The median is a poor predictor of the worst case. PaddleOCR’s p95 on SROIE (3,331.4 ms) is 11.2× its p50 (297.0 ms) — 5% of pages took over 3.3 seconds even though the median page took under 300 ms. For interactive workloads, this ratio, not the p50, decides whether a user waits 0.3 seconds or 3.3. docTR’s p95 (281.4 ms) stays tight at 2.6× its p50.
The mechanism is the first-page/prefill effect: GPU engines pay a one-time priming cost before steady inference, and batch scheduling can serialize slow samples. That is protocol-specific behavior — these p95 values reflect this benchmark’s run pattern (fixed test split, warm-then-scored), not a universal property of the engines. What the data shows is the shape of each engine’s tail under this protocol: PaddleOCR and Docling have long, heavy tails on SROIE; docTR, EasyOCR, and Unlimited-OCR keep theirs close to the median.
Source: computed from summary_metrics.csv — latency_p95_ms ÷ latency_p50_ms, sroie_2019 rows (ratio derived by division of the published values, e.g., 3331.3505 / 296.9897 = 11.22). PaddleOCR 11.2×, Docling 4.4×, docTR 2.6×, EasyOCR 2.3×, Tesseract 2.2×, Surya2 2.2×, PaddleOCR-VL 1.7×, Unlimited-OCR 1.6×.
Note what the ratio does not say: a tight p95/p50 ratio is not a speed claim. Unlimited-OCR has the tightest tail in the benchmark (1.6×) — but its p50 (1,600.7 ms) and p95 (2,521.9 ms) are both far slower than docTR’s p95 (281.4 ms). The ratio measures the shape of the distribution, not where the distribution sits. Read the two numbers together: a tight tail around a slow median is still slow.
Throughput Tells a Different Story: pages/min vs p50
Rank the engines by pages per minute and the ordering changes. docTR’s 449.3 pages/min vs Surya2’s 12.1 is a 37× spread — wider than the 24.5× p50 spread. But PaddleOCR, with an 11.2× p95 tail, still sustains 79.7 pages/min: close to EasyOCR’s 124.5, ahead of PaddleOCR-VL’s 68.2 and Docling’s 56.7. A slow median and a heavy tail do not prevent respectable batch throughput.
The reconciliation is the measurement basis: pages/min is wall-clock throughput including model init and scheduling, while p50 is steady-state single-page inference excluding load. The table below converts pages/min into the average wall-clock time per page it implies (60,000 ÷ pages/min — a derived estimate, not a measured figure) and shows the gap against p50. PaddleOCR’s implied wall-clock per page (752.7 ms) is 2.5× its steady-state p50 (297.0 ms); Surya2’s (4,940.0 ms) is 1.9× its p50 (2,668.0 ms). The overhead — init, scheduling, per-call pipeline costs — is invisible in p50 alone, which is why a low p50 does not automatically mean high throughput.
| Model (SROIE) | p50 (ms) | Pages/min | Wall-clock ms/page (derived) | Overhead vs p50 | Source |
|---|---|---|---|---|---|
| docTR | 108.7 | 449.3 | 133.5 | 1.2× | summary_metrics.csv · doctr/sroie_2019 row |
| EasyOCR | 413.6 | 124.5 | 481.8 | 1.2× | summary_metrics.csv · easyocr/sroie_2019 row |
| PaddleOCR | 297.0 | 79.7 | 752.7 | 2.5× | summary_metrics.csv · paddleocr/sroie_2019 row |
| Tesseract | 670.9 | 78.6 | 763.0 | 1.1× | summary_metrics.csv · tesseract/sroie_2019 row |
| PaddleOCR-VL | 694.3 | 68.2 | 880.3 | 1.3× | summary_metrics.csv · paddleocr_vl_vllm/sroie_2019 row |
| Docling | 732.0 | 56.7 | 1,059.1 | 1.4× | summary_metrics.csv · docling/sroie_2019 row |
| Unlimited-OCR | 1,600.7 | 34.4 | 1,742.5 | 1.1× | summary_metrics.csv · unlimited_ocr/sroie_2019 row |
| Surya2 | 2,668.0 | 12.1 | 4,940.0 | 1.9× | summary_metrics.csv · surya2/sroie_2019 row |
Table: summary_metrics.csv — latency_p50_ms / pages_per_minute, sroie_2019 rows. Wall-clock ms/page is a derived estimate (60,000 ÷ pages_per_minute, e.g., 60,000 / 79.7130 = 752.7) — arithmetic on the measured throughput, not a separately measured figure. Overhead = derived wall-clock ÷ measured p50. Pages/min is wall-clock including model init; p50 is steady-state excluding load.
A second reconciliation note on the extreme ends: Tesseract, the only CPU engine, sustains 78.6 pages/min on SROIE — essentially matching PaddleOCR’s 79.7 on GPU — with its p50 (670.9 ms, CPU) and wall-clock (763.0 ms) nearly identical, because a single-threaded CPU engine has little init overhead to hide. Its p95 (1,507.0 ms) is 2.2× its p50 — one of the tightest tails in the benchmark. The CPU-vs-GPU throughput parity is analyzed in detail in Tesseract vs PaddleOCR.
Interactive vs Batch: The 1,000 ms Threshold Reading
For a user waiting on a single page, 1,000 ms is a useful boundary — roughly the edge of what feels responsive. Read at p50, six of eight engines stay under it on SROIE: docTR (108.7 ms), PaddleOCR (297.0 ms), EasyOCR (413.6 ms), Tesseract (670.9 ms, CPU), PaddleOCR-VL (694.3 ms), Docling (732.0 ms). Only Unlimited-OCR (1,600.7 ms) and Surya2 (2,668.0 ms) cross it at the median.
Read at p95, the picture inverts: only two engines still fit under 1 second — docTR (281.4 ms) and EasyOCR (960.4 ms). Every other engine’s tail crosses it: PaddleOCR 3,331.4 ms, Docling 3,239.8 ms, Unlimited-OCR 2,521.9 ms, Tesseract 1,507.0 ms, PaddleOCR-VL 1,154.3 ms, Surya2 5,872.2 ms (summary_metrics.csv latency_p95_ms, sroie_2019 rows). If “interactive” means the worst case must feel responsive, the tail — not the median — is the selection criterion, and only two engines qualify.
Batch processing is the other regime, and it changes the economics in the engines’ favor. Model initialization is paid once per process/batch, so sequential or batched processing amortizes the init cost over more pages — per-page latency and per-page cost both fall as batch size grows. This is a derived argument from the measurement basis (pages/min includes init; p50 excludes it), not a new benchmark run — the same reasoning and its worked arithmetic appear in the cost page’s method.
Interactive-viable at the median
- docTR — 108.7 ms p50 / 281.4 ms p95
- PaddleOCR — 297.0 ms p50 / 3,331.4 ms p95
- EasyOCR — 413.6 ms p50 / 960.4 ms p95
- Tesseract (CPU) — 670.9 ms p50 / 1,507.0 ms p95
- PaddleOCR-VL — 694.3 ms p50 / 1,154.3 ms p95
- Docling — 732.0 ms p50 / 3,239.8 ms p95
p50 under 1,000 ms on SROIE 2019 (summary_metrics.csv latency_p50_ms, sroie_2019 rows). Medians feel fast; tails vary widely.
Interactive-viable at the tail
- docTR — 281.4 ms p95
- EasyOCR — 960.4 ms p95
p95 under 1,000 ms on SROIE 2019 (summary_metrics.csv latency_p95_ms, sroie_2019 rows). Only these two keep the worst case under one second.
Batch-only on receipts
- Unlimited-OCR — 1,600.7 ms p50 / 2,521.9 ms p95
- Surya2 — 2,668.0 ms p50 / 5,872.2 ms p95
- All engines at high volume — amortize init
p50 over 1,000 ms on SROIE (summary_metrics.csv latency_p50_ms, sroie_2019 rows). Batch or sequential processing amortizes init (derived from the measurement basis, not a new run).
CORD (Indonesian Receipts): Same Story, Different Tails
Swap the document set and the rankings roughly hold — docTR stays fastest and Surya2 stays slowest — but the language change reshapes the tails. On CORD v2, Surya2’s p95 blows up to 26,976.9 ms (≈27 s), 18.2× its own p50 (1,485.3 ms) and the benchmark’s extreme tail event — while its SROIE tail ratio was a modest 2.2×. Language content changes tail behavior; a VLM stable on English receipts can spike to half-minute waits on Indonesian ones.
CORD v2 is an Indonesian-language dataset with nested fields (menu, sub_total, total); none of the 8 engines was trained predominantly on Indonesian, so it doubles as a cross-language stress test. Its receipts are shorter and less text-dense than SROIE’s, which is why most engines get faster here — docTR drops to 100.4 ms p50 (500.4 pages/min), PaddleOCR to 108.2 ms p50 (141.0 pages/min). CORD is deliberately kept separate from the SROIE ranking in this benchmark (different language, different ground-truth structure); the point of this table is that latency moves with the document set and the tail can move dramatically.
Source: summary_metrics.csv — latency_p50_ms column, cord_v2 rows. docTR 100.4, PaddleOCR 108.2, EasyOCR 192.8, PaddleOCR-VL 249.3, Docling 338.0, Tesseract 474.6 (CPU-only), Unlimited-OCR 608.2, Surya2 1485.3. Steady-state latency, warm-then-scored (excludes model loading).
| Rank | Model | Type | p50 (ms) | p95 (ms) | p95/p50 | Pages/min | Source |
|---|---|---|---|---|---|---|---|
| 1 | docTR | Traditional OCR (GPU) | 100.4 | 235.9 | 2.3× | 500.4 | summary_metrics.csv · doctr/cord_v2 row |
| 2 | PaddleOCR | Traditional OCR (GPU) | 108.2 | 2,228.3 | 20.6× | 141.0 | summary_metrics.csv · paddleocr/cord_v2 row |
| 3 | EasyOCR | Traditional OCR (GPU) | 192.8 | 599.9 | 3.1× | 211.8 | summary_metrics.csv · easyocr/cord_v2 row |
| 4 | PaddleOCR-VL | Document-parsing VLM | 249.3 | 1,190.9 | 4.8× | 67.1 | summary_metrics.csv · paddleocr_vl_vllm/cord_v2 row |
| 5 | Docling | Pipeline parser | 338.0 | 1,333.2 | 3.9× | 123.2 | summary_metrics.csv · docling/cord_v2 row |
| 6 | Tesseract | Traditional OCR (CPU) | 474.6 | 1,011.3 | 2.1× | 108.9 | summary_metrics.csv · tesseract/cord_v2 row |
| 7 | Unlimited-OCR | Document-parsing VLM | 608.2 | 1,486.8 | 2.4× | 74.0 | summary_metrics.csv · unlimited_ocr/cord_v2 row |
| 8 | Surya2 | Document-parsing VLM | 1,485.3 | 26,976.9 | 18.2× | 11.2 | summary_metrics.csv · surya2/cord_v2 row |
Table: summary_metrics.csv — latency_p50_ms / latency_p95_ms / pages_per_minute, cord_v2 rows (100 samples each). p95/p50 ratios computed by division of the CSV values (e.g., 26976.9306 / 1485.3493 = 18.16). Tesseract ran CPU-only. Steady-state latency excludes model loading; pages/min is wall-clock including init.
Surya2’s CORD p95 of 26,976.9 ms is the single largest tail event in the benchmark — one page in twenty took about 27 seconds on a dataset where its median was 1.5 seconds. For context, that is 114× docTR’s CORD p95 (235.9 ms). Two engines that were near-identical on SROIE (Surya2 2.2×, Unlimited-OCR 1.6× tail ratios) diverge sharply on CORD (18.2× vs 2.4×) — a reminder that tail behavior is a property of the engine × document combination, not the engine alone.
Latency and Cost Run on the Same Clock
Because rented GPUs bill by wall-clock hour, latency and cost are the same measurement seen twice: docTR is both the fastest engine (108.7 ms p50) and the cheapest (cost_per_1000_pages 0.0479 on SROIE); Surya2 is both the slowest (2,668.0 ms p50) and the priciest (1.0609) — same split, same rate. The heavy VLMs are slow and pricey together; the lightweight traditional engines are fast and cheap together.
The correlation is not exact, because cost follows wall-clock pages/min (which includes init) while p50 excludes it — PaddleOCR’s p50 (297.0 ms) is faster than EasyOCR’s (413.6 ms), yet PaddleOCR’s per-1,000-page cost (0.2214) is roughly double EasyOCR’s (0.1098) because its wall-clock throughput is far lower (79.7 vs 124.5 pages/min; summary_metrics.csv cost_per_1000_pages / pages_per_minute, sroie_2019 rows). The full cost ranking, cost formula, and the arithmetic behind it live on the sibling OCR Cost per 1,000 Pages page — this page keeps its focus on latency and only borrows the correlation.
Frequently Asked Questions
Which is the fastest OCR engine per page?
docTR was the fastest in this benchmark: 108.7 ms p50 and 281.4 ms p95 per page on SROIE 2019 (summary_metrics.csv latency_p50_ms / latency_p95_ms, doctr/sroie_2019 row), sustaining 449.3 pages/min. The slowest measured, Surya2, was 24.5× slower at p50 (2,668.0 ms) and 37× slower on throughput (12.1 pages/min).
What is a typical OCR latency per page on a GPU?
Between 108.7 ms (docTR) and 2,668.0 ms (Surya2) p50 per page across 8 open-source engines on an RTX 4090, SROIE 2019 receipts (summary_metrics.csv latency_p50_ms, sroie_2019 rows) — a 24.5× spread on identical hardware. Steady-state, warm-then-scored, excluding model loading. The same engines on CORD v2 span 100.4 to 1,485.3 ms p50.
Why is OCR p95 latency so much higher than p50?
Because of the first-page/prefill effect and batch-scheduling behavior: GPU engines pay one-time priming costs and slow samples can be serialized, so the slowest 5% of pages — by definition of the 95th percentile — sit far from the median. PaddleOCR’s SROIE p95 (3,331.4 ms) is 11.2× its p50 (297.0 ms); docTR’s p95 (281.4 ms) is only 2.6× its p50 (summary_metrics.csv latency_p95_ms / latency_p50_ms, sroie_2019 rows). The p95 values are protocol-specific — they reflect this benchmark’s run pattern, not a universal engine property.
Why don’t OCR latency and pages per minute match?
Because they are different measurements. Pages/min is wall-clock throughput including model initialization and scheduling; p50 is steady-state per-page inference excluding model loading. On SROIE, PaddleOCR’s p50 (297.0 ms) implies roughly 200 pages/min at pure steady state, but measured throughput is 79.7 pages/min — a derived wall-clock time of 752.7 ms per page, 2.5× its p50 (summary_metrics.csv latency_p50_ms / pages_per_minute, paddleocr/sroie_2019 row). For batch jobs, plan around pages/min; for interactive waits, around p50 and p95.
What is a good OCR latency for interactive use?
Under 1,000 ms at the tail is the practical bar. On SROIE, six of eight engines stay under 1 second at p50 (docTR 108.7, PaddleOCR 297.0, EasyOCR 413.6, Tesseract 670.9, PaddleOCR-VL 694.3, Docling 732.0 ms), but only docTR (281.4 ms) and EasyOCR (960.4 ms) keep their p95 under 1 second (summary_metrics.csv latency_p50_ms / latency_p95_ms, sroie_2019 rows). If the worst case must feel responsive, the p95 decides; if the typical case is enough, the p50 decides.
Why did one OCR engine take 27 seconds on a receipt?
Surya2’s p95 on CORD v2 was 26,976.9 ms (≈27 s) — 18.2× its own p50 (1,485.3 ms) and the largest tail event in the benchmark (summary_metrics.csv latency_p95_ms / latency_p50_ms, surya2/cord_v2 row). Its tail on SROIE was a modest 2.2×, so the spike is specific to the engine × CORD combination — language content and scheduling changed the tail behavior, not just the median.
Is Tesseract OCR fast enough without a GPU?
On CPU, Tesseract sustained 78.6 pages/min on SROIE 2019 with a p50 of 670.9 ms — slower than the GPU traditional engines but faster than all three VLMs on throughput (PaddleOCR-VL 68.2, Docling 56.7, Unlimited-OCR 34.4, Surya2 12.1 pages/min) and with a tight tail (2.2×). It is CPU-only in this benchmark (summary_metrics.csv tesseract/sroie_2019 row); whether it is “fast enough” depends on your volume and whether you need fields rather than text — see the CPU cost profile in Tesseract vs PaddleOCR.
Does OCR get faster per page as you process more pages?
Yes, up to a steady-state floor — as a derived effect of the measurement basis, not a new run. Model initialization is paid once per process/batch (it is inside pages/min but outside p50), so batch or sequential processing amortizes it: at high volume, per-page wall-clock time approaches the steady-state p50 plus scheduling. docTR’s SROIE numbers are already close to that floor (p50 108.7 ms vs derived wall-clock 133.5 ms per page); Surya2’s (p50 2,668.0 vs 4,940.0 ms) shows more init overhead to amortize.
Where do the latency numbers on this page come from?
Every figure is a row of the first-party benchmark’s published results/summary_metrics.csv (latency_p50_ms, latency_p95_ms, pages_per_minute) hosted at ImageToTableai/benchmark-ocr, with one redacted manifest.json per run recording the measurement mode (warm_then_scored), model versions, and environment fingerprints. Dataset definitions come from the SROIE 2019 and CORD papers cited below.
Methodology & Sources
Protocol
This page reports the latency dimension of an independent, reproducible benchmark run (official tier) — not a survey of third-party claims. Fixed test splits only: SROIE 2019 test (361 English receipts, flat fields company/date/address/total) and CORD v2 test (100 Indonesian receipts, nested fields menu/sub_total/total); training splits were never evaluated. Every (engine × dataset) pair reused the same images, the same ground truth, and the same measurement protocol (warm_then_scored: a fixed warm-up pass precedes the scored pass). All 16 runs completed with error_rate 0.0 (summary_metrics.csv error_rate column). Data was collected in August 2026.
Runtime Environment
- Hardware: all GPU runs on a single NVIDIA RTX 4090 (24 GB). Tesseract ran CPU-only (compute_type=cpu) and is labeled as such on every table. One GPU tier, one location/time — the range statement on this page.
- Engines: all models run out-of-the-box, no fine-tuning. Versions locked per the run manifests: Tesseract 5.3.4, PaddleOCR 3.7.0, EasyOCR 1.7.2, docTR v1.0.1, Docling 2.119.0, Surya2 0.22.1, Unlimited-OCR vLLM-served, PaddleOCR-VL 1.6.
- Measurement mode:
warm_then_scored— latency figures are steady-state per-page inference excluding model loading; pages/min is wall-clock including model init (per run’sperformance.run_wall_time_msin the redacted manifests). - Run tier: all 16 runs are
official; only official runs are eligible for publication per the benchmark protocol (reports/receipt_v1_official_protocol.md).
Metric Definitions
- p50 (median latency): the median per-page inference time in steady state — 50% of pages were faster. Excludes model loading (warm-then-scored).
- p95 (tail latency): the 95th percentile per-page inference time — 5% of pages took longer. Reflects first-page/prefill effects and batch-scheduling behavior under this run’s protocol; protocol-specific, not a universal engine constant.
- p95/p50 ratio: derived by division of the two CSV columns (e.g., 3331.3505 / 296.9897 = 11.22). A shape indicator of the tail, not a speed claim.
- Pages per minute: wall-clock throughput including model init. The batch-jobs and billing view.
- Wall-clock ms/page (derived): 60,000 ÷ pages/min — arithmetic on the measured throughput, labeled as a derived estimate, not a measured figure.
Source List
- summary_metrics.csv (GitHub raw). 16 rows = 8 engines × 2 receipt datasets (sroie_2019, cord_v2). Columns include latency_p50_ms, latency_p95_ms, pages_per_minute, compute_type, error_rate. Every latency and throughput figure on this page traces to a row here.
- ImageToTableai/benchmark-ocr repository. Public repo hosting the result CSVs, redacted run manifests, frozen protocol (
reports/receipt_v1_official_protocol.md), and dataset sample lists (fixed test splits) for reproduction. - results/manifests/ (GitHub). One redacted manifest.json per published run (16 runs) with the environment fingerprint, model versions, measurement mode (
warm_then_scored), wall-clock time, and artifact hashes. - Huang et al., "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2019). SROIE 2019 dataset definition, task structure, and license (CC-BY-4.0).
- Park et al., "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing" (2020). CORD v2 dataset definition, nested field schema, and license (CC-BY-4.0).
Limitations
- Single GPU tier, single location/time: all GPU figures come from one RTX 4090 at one site, collected August 2026. Other GPUs, multi-GPU serving, and different scheduling shift latency and throughput; measure on your own hardware before committing.
- Receipts only: SROIE (English) and CORD (Indonesian) receipts. Latency on invoices, forms, contracts, or long documents is unmeasured; the rankings above are not generalizable beyond receipts — and CORD’s results are not merged into the SROIE ranking.
- p95 is protocol-specific: p95 values reflect first-page/prefill effects and batch-scheduling under this benchmark’s run pattern (fixed 361/100-page splits, warm-then-scored). They are a shape reading of this run, not a universal worst-case guarantee.
- p50/p95 exclude model loading, pages/min includes it: the two bases are deliberately different (steady-state inference vs wall-clock throughput) and are never presented as the same number. Derived wall-clock per page (60,000 ÷ pages/min) is arithmetic on measured throughput, not a separately measured figure.
- CPU/GPU asymmetry: Tesseract (CPU) is compared against GPU-accelerated engines; its latency reflects CPU hardware. It is marked on every table but the asymmetry is inherent to the comparison.
- No cloud/API models: AWS Textract, Google Document AI, Azure AI Document Intelligence, and hosted OCR/VLM APIs are not included; their latency models (network, per-call metering, autoscaling) differ fundamentally from the local engines measured here.
- No batch-size sweep: init amortization is derived from the measurement basis (pages/min includes init, p50 excludes it) and cross-referenced to the cost page’s method — no batch-size experiment was run, so per-batch scaling is an estimate, not a measurement.
- Sample size and version pinning: 361 + 100 samples; results hold for the August 2026 model versions listed above. Newer engine releases may shift latency; single-digit-percent differences should be treated as noise.
Related references: OCR Cost per 1,000 Pages · OCR pipelines against document-parsing VLMs · docTR vs Surya2: CER Tie, Cost Gap · docTR vs Docling: Single-Pass vs Pipeline · Tesseract vs PaddleOCR: Legacy CPU vs Modern GPU
Related reading: what AI OCR accuracy actually measures · AI extraction from images versus traditional OCR pipelines · AI Document Extraction Pricing (2026)