Exception Rates in Document Automation:
Queue, Cost & Threshold (2026)
Last reviewed: 2026-08-10 · Data coverage: Partial · Sources: 18 independent studies
What this page does NOT cover: Manual data entry error rates at the keystroke/field level (see Manual Data Entry Error Rates), per-invoice processing costs (see Document Processing Cost per Record), OCR character-level accuracy (see OCR Accuracy by Document Type), or vendor-specific performance claims presented as independent facts.
All numbers below are based on publicly available third-party studies and industry reports cited in the methodology section. Where sources disagree, conservative ranges (lowest–highest reported values) are used. Vendor-reported figures are explicitly labeled as such. This page is a directional benchmark, not a precise forecast for any single organization.
Consider a team processing 2,000 invoices a month. Ardent Partners' 2024 survey (n=212) puts the average invoice exception rate at 14% — that is 280 invoices pulled out of the automated flow for special handling, each consuming 15–45 minutes of human work (IOFM 2024) on top of normal processing. Meanwhile only 32.6% of invoices run straight through with no human touch at all. The gap between the two numbers — 14% "exceptions" versus 67% "some human contact" — is not a measurement error: it is the difference between the special-handling queue that burns hours and the normal approval touches that don't. And the spread between teams is the real story: best-in-class teams hold exceptions to 9%, while everyone else sits at 22%.
Exception rate is the hidden operational tax on every document automation deployment. This page opens with the AP scenario because it has the best data, then stacks on the dimensions that change the number: automation type, root cause, industry, and finally the confidence threshold that sets the boundary.
The Exception Queue: What 14% Actually Looks Like
An exception is a document that fails validation or matching and is routed out of the automated flow for special handling — a missing purchase order number, a price that doesn't match the PO, a duplicate submission, a field the extraction engine couldn't read with confidence. It is distinct from the broader set of documents that receive any human touch: an invoice can go through an approval workflow (a normal, cheap touch) and never be an exception. IOFM's benchmark puts the manual cost of a typical invoice at 12.5 minutes, while an exception adds 15–45 minutes on top (IOFM 2024); under traditional automation, each exception takes 45–60 minutes to investigate and resolve (APQC 2025, via Peakflo).
The cost of the queue is disproportionate to its size: exceptions typically represent only 5–15% of document volume but 30–50% of total processing cost (Artificio 2025), because every exception consumes the most expensive resource in the workflow — skilled human attention. At 14% of 2,000 invoices with a 30-minute midpoint, that is 140 hours per month — nearly one full-time employee — doing nothing but resolving exceptions.
Breakdown by Automation Type
The exception rate is not a property of the documents — it is a property of the automation generation that processes them. Template-based capture tops out around 10–20% STP because every layout variation falls back to a human. ML-based IDP platforms with validation rules push the ceiling higher. The independent survey anchors — Ardent's 32.6% average and 49.2% best-in-class, Hackett's 60% among adopters — describe real organizations; the 85–92% figures for agentic platforms come from vendor production reports and should be read as upper bounds, not industry averages.
Sources: Ardent Partners (2025) — 32.6% average / 49.2% best-in-class, n=212; The Hackett Group (2025) — 60% among AP-solution adopters; Hypatos (2019) — template-era 10–20% STP (industry claim); Hypatos (2026) — 85–92% agentic STP, vendor-reported production figures. Conservative midpoints used where sources disagree; vendor figures labeled as such.
| Automation Type | Typical STP Rate | Implied Exception Rate | What Sets the Ceiling | Source | Year |
|---|---|---|---|---|---|
| Manual entry | ~0% (by definition) | 100% touch | Every document is keyed by hand | Definitional | — |
| Template / zonal OCR | 10–20% | 80–90% | Layout change = re-template or fallback | Hypatos (industry claim, D-level) | 2019 |
| Automation adopter (survey average) | 32.6% | ~14% exception rate | Mixed document intake, partial automation | Ardent Partners | 2024/2025 |
| Best-in-class | 49.2% | 9% | Supplier e-invoicing enablement + validation rules | Ardent Partners; Hackett (60% adopters) | 2024/2025 |
| Agentic AI platforms | 85–92% | 8–15% | Downstream matching + exception resolution automated | Hypatos production claims (vendor-reported) | 2026 |
The exception-rate column for survey rows is the org-level rate (Ardent: 14% average, 9% BIC); for template/agentic rows it is the implied inverse of the vendor-reported STP figure. The 30–40% exception rate reported for organizations with traditional automation but without intelligent exception handling (APQC 2025, via Peakflo) sits between the "template" and "adopter" rows — the automation captures the document but hands the reconciliation problem back to humans.
Breakdown by Root Cause
The uncomfortable truth about exception queues is that most organizations cannot decompose their own: industry practice lumps everything under "data quality," which is too vague to fix. The sources below agree on the ordering — matching failures dominate, extraction errors are a distant second — and offer a few quantified anchors. Three-way match failures are the single most common cause (Zamp 2026; Transcepta 2024), and the root of that root is often upstream: over 30% of purchase-order discrepancies trace to manual data entry (Resolvepay 2025; Transcepta), extending payment cycles by 7–10 business days per disputed invoice (Transcepta).
| Root Cause | Frequency / Evidence | Source | Year |
|---|---|---|---|
| PO / three-way match failures | Most common exception cause; >30% of PO discrepancies stem from manual entry errors; extends payment cycles 7–10 business days | Zamp; Transcepta; Resolvepay | 2024–2026 |
| Missing / malformed reference data | Missing PO numbers or invalid vendor codes on otherwise valid invoices; common with smaller or international suppliers | Zamp; Transcepta | 2024–2026 |
| Extraction / OCR errors | Misread characters or fields route to review; a common cause per Transcepta's exception taxonomy | Transcepta | 2024 |
| Duplicate submissions | 0.1–3% of total invoice volume; at $500M annual invoice spend, weak detection can mean $1.5M paid twice | Zamp (industry consensus value) | 2025/2026 |
| GL coding / approval stalls | Coding mismatches and slow approvals are structural causes; 60–70% of AP team time spent on exception handling vs strategic work | APQC 2025 via Peakflo; Transcepta | 2024/2025 |
Root-cause share percentages are rarely published — Transcepta, Zamp, and APQC provide the ordering and the few quantified anchors above (PO discrepancy sources, duplicate rates, exception-handling time share). The 60–70% team-time figure is an AP-specific APQC benchmark. See Limitations for what could not be found on this dimension.
Breakdown by Industry
The exception question generalizes far beyond AP — but the data quality does not. AP is the best-measured domain; healthcare claims and insurance are covered by payer-industry estimates that are directionally consistent. The pattern is the same everywhere: roughly a fifth to a third of documents still need human handling, and the highest-cost documents are the ones most likely to be flagged.
| Industry / Workflow | Exception / Manual-Review Rate | Straight-Through Rate | Source | Year |
|---|---|---|---|---|
| AP / invoice processing (average) | 14% exception rate | 32.6% | Ardent Partners (n=212) | 2024/2025 |
| AP / invoice processing (best-in-class) | 9% | 49.2% | Ardent Partners | 2024/2025 |
| Healthcare claims (payer side) | ~20% require manual inspection (the highest-cost, most complex claims) | ~80% auto-adjudicated | Policy brief citing CAQH / Larsson (2017) | 2017/2026 |
| Insurance claims (P&C, low-severity) | — | 5–10% manual baseline → 20–40% with automation; leading carriers 40–60% for auto-glass type claims | Peakflo claims KPI research | 2025/2026 |
| Insurance — duplicate payments | 2–5% of payments (manual) → <0.5% automated | — | Peakflo claims KPI research | 2025/2026 |
Healthcare claims figures are payer-industry estimates (Larsson 2017, cited via policy research) — labeled by year as the underlying data predates the 3-year window. Insurance STP ranges are industry KPI tables compiled from multiple unnamed sources (Peakflo), presented as directional. No independently audited exception-rate benchmarks were found for logistics, legal, or HR document workflows (see Limitations).
Confidence Thresholds: The Lever That Sets the Queue Size
Exception rates are not only a property of the model — they are a dial. Every extraction engine assigns a confidence score to each field, and the threshold you set decides which documents flow through and which land in the review queue. Lower the threshold and more documents automate, but more errors pass through uncaught. Raise it and the queue grows, catching more errors but burning reviewer hours on documents that were fine. The four metrics that matter are the four sides of this dial:
| Metric | Definition | What Drives It | Cost of Getting It Wrong |
|---|---|---|---|
| Exception rate | % of documents routed to special human handling | Validation failures, low confidence, business rules | Queue labor + cycle-time delay |
| Straight-through rate (STP) | % of documents processed with zero human touch | Inverse of the queue, minus approval-only touches | Low automation = high cost per document |
| False positive rate (FP) | Documents flagged for review that are actually correct | Threshold set too high (over-caution) | Wasted reviewer hours — the hidden cost of "safe" thresholds |
| False negative rate (FN) | Errors that slip through without any review | Threshold set too low (over-automation) | Downstream errors — an extracted wrong amount posted to the ERP |
Peer-reviewed research on selective prediction formalizes this tradeoff: "silently wrong is more dangerous than visibly absent" — an LLM extraction that is wrong but confident does more damage than one that honestly flags itself for review (arXiv 2606.24420). On the DocILE invoice benchmark, where LLM extraction has a 26% natural failure rate, a multi-signal confidence engine that automates the top 80% of documents by confidence reaches 99.1% accuracy — versus 73.3% if everything is automated blindly (Kumar 2026, arXiv:2606.24420). In other words: correctly identifying which 20% to review is worth more than improving the model's raw accuracy.
The engineering literature proposes a double-threshold policy for exactly this problem: a lower bound (auto-reject / route to review) and an upper bound (auto-accept), with the ambiguous band between them going to human reviewers. The optimization is explicitly framed as balancing accuracy against the cost of human review (Muric & Minton 2026, arXiv:2601.05974). Applied to invoice workflows, this means the right threshold is not a fixed number but a function of review capacity and error tolerance: a team with reviewer hours to spare should run a more conservative threshold; a team drowning in queue volume with a tolerant downstream system can afford a more aggressive one. A 2025 evaluation of invoice extraction methods makes the same point in reverse — proposing STP rate, manual review rate, and false auto-approval rate as the three business-level KPIs, with per-field precision-recall threshold calibration as the tuning mechanism (Yashwant et al. 2025, arXiv:2510.15727).
The asymmetry of the two error types has been documented since the earliest OCR economics research: false positives (words the system claims it read but read wrong) are rare; false negatives (words it simply failed to read) are far more common (Hartley & Crumpton 1999, arXiv:cs/9902009). The correction economics follow: an over-cautious threshold that flags correct extractions (FP) burns reviewer time but is cheap; an over-aggressive threshold that lets wrong extractions through (FN) silently corrupts downstream systems — exactly the failure mode the 2026 confidence research warns about.
Impact by Role
Who the exception queue falls on changes what the benchmark means: for clerks it is hours of rework, for managers it is a leading indicator that predicts cycle time, for executives it is the difference between best-in-class cost and everyone else's.
Sources: IOFM (2024) — exception add-on time; Ardent Partners (2025) — exception rate, STP, top-challenge ranking, cost; The Hackett Group (2025) — adopter touchless rate; APQC 2025 via Peakflo — per-exception investigation time. The ~140 hours/month figure is arithmetic from cited rates, not a published survey number.
How to Use This Data
You don't need a full audit to size your exception queue. Four numbers — volume, exception rate, minutes per exception, and burdened hourly cost — turn "exceptions are a problem" into a headcount and dollar figure.
- Pick a workflow and count its volume — for example, 2,000 invoices/month through AP.
- Apply the exception-rate range from this page. For day-to-day planning use the average of 14% (Ardent 2025) or the conservative 9–22% band (best-in-class vs other firms). If your organization runs traditional automation without intelligent exception handling, APQC's 30–40% range is the more honest planning number.
- Multiply by time and cost per exception. At 14% of 2,000 invoices = 280 exceptions/month × 30 min (IOFM midpoint) = 140 hours ≈ 0.9 FTE; at a fully-burdened $45/hour, that is roughly $6,300/month of exception-handling labor. The average cost to rectify a single manual exception often exceeds $50 once research and follow-up are included (Transcepta 2024).
- Compare to the cost of moving the dial. If your exception rate is 22% (other-firm average) and you could reach best-in-class 9%, that is 260 fewer exceptions per 2,000 invoices — 130 hours and ~$5,800/month saved before counting faster cycle time, fewer late payments, and recovered early-payment discounts. For the cost side of the same math, see Document Processing Cost per Record.
These are rough-estimate formulas, not financial advice. Your exact numbers will vary by industry, document mix, and existing processes. The goal is to move from "we don't know" to "here's a defensible range."
Frequently Asked Questions
What is a normal invoice exception rate?
The average invoice exception rate is 14% (Ardent Partners 2025, n=212) — best-in-class teams hold it at 9%, while all other firms average 22%. Historically the average was higher (~22.5–23.2% in earlier Ardent studies, per Transcepta and ARDEM), so the 14% figure reflects a downward trend as automation matures. Exception rates above 15% are widely treated as a signal of systemic process failures (Transcepta).
What is straight-through processing (STP) rate?
STP rate is the percentage of documents processed end-to-end with zero human intervention — from receipt to posting/payment, no clicks, no corrections, no approvals. The AP average is 32.6%, best-in-class reaches 49.2% (Ardent 2025), and AP-solution adopters average 60% (Hackett 2025). It is the inverse concept of the exception rate, but the two are not exact mirror images: a document can pass through an approval workflow (human touch, not an exception) and still not count toward STP.
What is the difference between exception rate and STP rate?
Exception rate counts documents that fail validation or matching and need special handling; STP rate counts documents that need zero human contact. They are related but not complements: only ~67% of invoices have any human touch (100% − 32.6% STP), but only 14% are exceptions — the other ~53% of touched invoices go through routine approval or verification without special handling. Exception rate is the narrower, more expensive queue; STP rate is the broader automation ceiling.
How much does an invoice exception cost?
Each exception adds 15–45 minutes of human handling (IOFM 2024) — 45–60 minutes under traditional automation (APQC 2025) — and the average cost to rectify a single manual exception often exceeds $50 (Transcepta 2024). At scale this is material: exceptions are only 5–15% of document volume but account for 30–50% of total processing cost (Artificio 2025).
What causes invoice exceptions?
Purchase-order and three-way match failures are the single most common cause (Zamp 2026; Transcepta 2024), with over 30% of PO discrepancies tracing to manual data entry (Resolvepay; Transcepta). Other structural causes: missing or invalid PO/vendor reference data, extraction/OCR errors, duplicate submissions (0.1–3% of volume), and GL coding or approval stalls (Zamp; Transcepta).
How do confidence thresholds affect exception rates?
Confidence thresholds are the dial that sets the queue size: raise the threshold and more documents route to review (higher exception rate, fewer errors slipping through); lower it and more automate (lower exception rate, more false negatives). On the DocILE invoice benchmark, routing only the lowest-confidence 20% of documents to review yields 99.1% accuracy on the automated 80% — versus 73.3% if everything is automated blindly (arXiv 2606.24420). The optimal threshold balances review cost against error tolerance (arXiv 2601.05974).
What is a good straight-through processing rate?
For AP, 32.6% is the cross-industry average and 49.2% is best-in-class (Ardent 2025); organizations that have adopted AP automation solutions average 60% (Hackett 2025). Vendor-reported production rates for agentic platforms run 85–92% but are not industry averages — treat them as upper bounds achievable under controlled conditions (Hypatos 2026, vendor-reported).
How much time does an exception add to processing?
An exception adds 15–45 minutes of human handling beyond the 12.5-minute average invoice touch (IOFM 2024), and resolution still takes 5.0 days of calendar time even at the fastest quartile of organizations (APQC exception-resolution measure, n=461). Under traditional automation the per-exception investigation alone runs 45–60 minutes (APQC 2025, via Peakflo).
Methodology & Sources
How This Page Was Built
This page aggregates data from 14 independent studies and reports spanning 1999–2026. Sources were selected based on three criteria: (1) publicly documented methodology with stated sample size where available, (2) data collected within the last 3 years where possible (older sources — Hartley 1999, Larsson 2017, Hypatos 2019 — are explicitly labeled by year), and (3) relevance to business document processing workflows rather than academic OCR research alone. Where multiple sources report the same dimension, conservative ranges (lowest–highest reported values) are used. Where sources disagree significantly (e.g., 14% vs 22.5% exception rate), the disagreement is explained by measurement year rather than averaged away. Vendor-reported STP figures (Hypatos, Conexiom-type claims) are never used as independent facts — they appear only in the automation-type dimension, explicitly labeled.
Source List
- Ardent Partners — AP Metrics That Matter in 2025 (eBook based on State of ePayables 2024). Survey of 212 AP professionals. Provides the exception rate 14% average (9% BIC vs 22% others), STP 32.6% average / 49.2% BIC, 53% citing exceptions as top challenge, cost per invoice $9.40 average / $2.78 BIC.
- Ardent Partners — State of ePayables 2023. Industry survey. Provides the 2023 STP baseline (32.4%) and BIC 2.1x STP advantage used for cross-year context.
- APQC — Process Benchmarking 2024 (via Nexus AP research). Nonprofit benchmarking organization. Provides AI auto-resolution of 60–70% of exceptions; exception-resolution cycle time 5.0 days (25th percentile, n=461, from APQC OSB). APQC's official measures pages are bot-protected; figures cross-checked via Nexus AP and Peakflo transcriptions.
- IOFM — AP Benchmarking / Ask the Expert (2024). Institute of Finance & Management, the leading AP association. Provides the 12.5-min manual invoice touch time and +15–45 min per exception. Primary report paywalled; figures transcribed via Nexus AP and cross-checked against Levvel's ranges.
- Transcepta — The Top 7 Causes of Invoice Exceptions (2024). AP network/industry education content. Provides the best-in-class <9–10% exception-rate benchmark, >15% systemic-failure signal, $50+ per-exception rectification cost, >30% of PO discrepancies from human error, and 7–10 day cycle extensions.
- The Hackett Group — AP Solutions Research (2025). Consulting/benchmark firm. Provides the 60% average touchless rate among AP-solution adopters.
- Levvel Research (now Deloitte) via Ascend Software (2025). Industry analyst research roundup. Provides ~30% industry touchless vs 60–80% high performers; manual invoice error ~2% vs <0.8% automated (IOFM-sourced).
- Hypatos — Document AI (2019) and Business Case for Agentic AI in Finance Ops (2026). Vendor research content (D-level). Provides template-era 10–20% STP claim and 85–92% agentic production STP — used only as labeled vendor-reported figures for the automation-type dimension.
- Zamp Finance — Root Causes of AP Invoice Exceptions (2025/2026). Vendor education content (D-level framework). Provides the four root-cause taxonomy (three-way match, reference data, duplicates at 0.1–3% of volume, approval stalls) and the 3–4% above-contract pricing slip-through estimate.
- Resolvepay — Statistics on PO-Mismatch Delays (2025). Industry media. Provides the >30% share of PO discrepancies from manual entry and the $25–50 error-correction cost range.
- One Percent Steps policy brief — Real-Time Adjudication (2026, citing CAQH and Larsson 2017). Policy research. Provides the ~80% auto-adjudication / ~20% manual-review split for health claims and the manual-review cost premium.
- Peakflo — AI Agents & Exception Management (2025/2026, citing APQC 2025) and Claims Automation KPIs (2026). Vendor research roundups. Provides APQC 2025 figures (30–40% exceptions under traditional automation, 45–60 min per exception) and insurance STP/duplicate-payment KPI ranges.
- Artificio — Handling Complex IDP Exceptions (2025). Vendor education content. Provides the exceptions-5–15%-of-volume / 30–50%-of-cost ratio and the healthcare exception-review case (45 → 18 min).
- Kumar, N. — Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction (arXiv:2606.24420, 2026). Peer-reviewed-track academic paper. Provides the DocILE 26% natural failure rate, 80%-coverage 99.1% accuracy vs 73.3% base rate, and the "silently wrong" formulation of the FP/FN tradeoff.
- Muric, G. & Minton, S. — A Framework for Optimizing Human-Machine Interaction in Classification Systems (arXiv:2601.05974, 2026). Academic paper. Provides the double-threshold policy framework balancing accuracy against human review cost.
- Yashwant, S. et al. — Invoice Information Extraction: Methods and Performance Evaluation (arXiv:2510.15727, 2025). Academic paper. Provides the STP / manual-review / false-auto-approval KPI framework and per-field threshold calibration.
- Hartley, R. & Crumpton, K. — Quality of OCR for Degraded Text Images (arXiv:cs/9902009, 1999). Academic paper. Provides the classic FP-rare / FN-common asymmetry in OCR correction economics — a methodology constant, labeled by year.
- ARDEM — True Cost of Invoice Exceptions (2024, citing Ardent Partners). BPO education content. Provides the historical ~23.2% exception-rate figure and the 15-min-per-exception / 12-FTE-for-500k-invoices arithmetic used for trend context.
Limitations
- Geographic coverage: The majority of sources are US/North America-based (Ardent, APQC, IOFM, Hackett, Transcepta, Resolvepay). No reliable published exception-rate benchmarks were found for APAC, Latin America, or emerging markets; the healthcare claims figure is US-specific (CAQH).
- Methodology constraints: Several key figures (IOFM exception add-on, APQC auto-resolution rate) come from paywalled or bot-protected primary reports transcribed by third parties — the secondary route is stated next to each figure. AP survey figures are self-reported, not measured. Root-cause shares are rarely published; the taxonomy relies on vendor education content (Transcepta, Zamp) whose ordering is corroborated across two independent sources but whose exact percentages are not.
- Temporal gaps: The most recent primary AP data is 2024–2025 (Ardent, IOFM). The 14% exception rate is the latest point in a series that began around 22.5–23.2% (labeled historical). Healthcare claims data dates to Larsson 2017. No reliable year-over-year exception-rate series exists outside Ardent.
- What we could not find: (1) No independently audited false positive / false negative rates for production document automation exist in published literature — the FP/FN distinction is supported by academic coverage-accuracy data (arXiv 2606.24420) and conceptual frameworks, not by industry benchmarks; (2) no citable exception-rate research from Kofax/Tungsten was found — their published content is product-capability marketing without methodology; (3) no published exception benchmarks for logistics, legal, or HR document workflows; (4) no independent (non-vendor) measurements of agentic-AI STP rates. These gaps are real; we state them rather than fill them with marketing numbers.
Related references: Manual Data Entry Error Rates · Invoice Processing Time Benchmarks · Document Processing Cost per Record
Related reading: How accurate is AI document extraction? · Scaling invoice processing without headcount