Batch Data Extraction

Batch Data Extraction: Process Multiple Documents into One Structured Spreadsheet Without Per-Format Setup

Manually extracting data from a stack of documents takes roughly 3 minutes per page — batch extraction processes each page in 5-10 seconds, from any format, into a single spreadsheet. No templates, no pre-sorting.

5-10s per page · Any format (PDF/JPG/PNG) · Mixed batch · Up to 99% accuracy on printed text

Batch Process
PDF / JPG / PNG
One Spreadsheet
Named Columns

What You Can Extract from a Batch of Documents

Name the columns you need once — the AI extracts those values from every document in your batch by understanding what they mean, regardless of format or layout. Each document becomes one row in a single spreadsheet.

Document Date
Document Number / ID
Vendor / Sender Name
Total Amount
Tax Amount
Currency
Line Item Description
Quantity
Unit Price
Line Total
Due Date
Category (AI Inferred)

You type these column names once. The AI applies them to every document in the batch — PDFs, scans, and photos alike.

Batch Data Extraction Sounds Simple — Until You Actually Try to Set It Up

Every tool promises "upload multiple files, get one spreadsheet." Most deliver that only if every document shares the same template and layout. When vendors differ, formats mix, and source quality varies — the real world — the setup cost reveals itself.

Where Traditional Batch Tools Add Hidden Setup Costs

01

One template per document layout means one template per vendor. Template-based tools demand a parser per layout — when every vendor's invoice is positioned differently, you build and maintain a separate configuration for each one. Even AI-powered tools like Nanonets or Rossum require per-document-type training or field mapping before they can process a batch. A user on Reddit described the situation bluntly: "We're manually copying data from PDFs into Excel every week and it's taking so much." The template solution doesn't eliminate the work — it shifts it from copying to configuring.

02

Mixed formats don't batch together in most tools. Native PDF, scanned image, and mobile photo require different processing pipelines in traditional OCR tools. Users end up running separate batches per format, then manually merging the outputs. One developer on Reddit put it starkly: "PDF is a human readable output format. It isn't a computer readable input format" — format-specific pipelines will always struggle with mixed batches.

03

Batch size and progress visibility are afterthoughts. Many tools process files one at a time behind a single progress bar. When one document fails, there's no way to know which one or why. The real bottleneck isn't AI speed — it's the lack of per-document status that forces re-verifying the entire batch manually.

How Column-Name Extraction Eliminates Setup Before the First Batch

01

You define the output, not the input rules. Instead of building a parser for each layout, you type the column names you need — Document Date, Vendor, Total Amount, Tax, Due Date — and the AI finds those values on every document by understanding what they mean. "Invoice Date" on one vendor's PDF and "Transaction Date" on a scanned receipt from another all resolve to your "Document Date" column. The column names themselves are the only configuration, and they apply to every document in the batch regardless of vendor, format, or layout. This is Custom Column Extraction: the AI reads for semantic meaning rather than enforcing positional rules.

02

All formats share one pipeline — no pre-sorting required. Drop a folder of PDF invoices, scanned JPG receipts, and mobile PNG photos into the same upload. The visual language model reads every file the same way: by looking at the pixels and extracting the named columns. No "PDF pipeline" vs "image pipeline" — just one batch, one column definition, one output spreadsheet.

03

Per-document progress and independent processing — one bad file doesn't sink the batch. Each document is processed independently with its own status tracked in real time. The batch view shows exactly which documents have completed, which are processing, and which encountered issues — no single progress bar hiding individual failures. Processing takes 5-10 seconds per page (vs ~3 minutes manual entry per page), and the output is one consolidated XLSX or CSV file with the columns you defined.

Batch Data Extraction in Practice: From a Folder of Mixed Files to One Clean Spreadsheet

Here's what the workflow looks like when you skip the per-format setup and go straight to defining your output.

1

Upload the Entire Batch — Formats Don't Matter

Your month-end folder contains 30 files: 18 PDF invoices from 10 different ERP systems, 7 scanned receipts as JPGs from field offices, and 5 mobile photos of handwritten delivery notes. Drag all 30 into the upload area. PDF, JPG, PNG, WebP — no pre-sorting, no "select document type" step. Batch upload handles roughly 30 files at once, each up to 10 MB.

2

Name Your Columns Once — They Apply to Every Document

Type Document Date, Vendor, Invoice Number, Total Amount, Tax Amount, Due Date. These become the exact headers in your output spreadsheet. The AI maps "Invoice Date" on digital PDFs, "Transaction Date" on scanned receipts, and "Date" on handwritten notes all to your "Document Date" column — by understanding what each field means, not by matching a document-type template. Optionally add a computed column like Line Total (Qty × Unit Price) and the AI performs the arithmetic during extraction.

3

One Spreadsheet Outputs — Every Document Is a Row

Processing takes 5-10 seconds per page. The output is a single XLSX file: 30 rows, one per document, with the columns you defined. If a vendor's invoice doesn't show tax, that cell is empty — not a fabricated zero, not a batch failure. Each row is independently extracted, so one poor-quality scan doesn't affect the other 29. The AI also supports Inferred Columns — define a "Category" column with options like "Invoice / Receipt / Delivery Note / PO" and the AI classifies each document as it extracts, producing both the data and the document-type label in one pass.

When Batch Extraction Works Best — and When to Be Cautious

Batch extraction reliably handles mixed formats and varied layouts when the target fields are labeled. Understanding the boundaries keeps your batches running smoothly.

When It Works Best

Mixed-format batches sharing common field concepts. When every document carries a date, total, and vendor name at different positions, the AI maps them to your named columns without per-document configuration.

Clear printed labels near values. Up to 99% accuracy when fields are labeled — "Invoice Date:" and "Total:" at any page position.

~30 documents per batch with per-file status visibility. Each document's status (processing / completed / error) visible in real time. Larger volumes split into multiple batches, each producing an independent spreadsheet.

When to Be Cautious

Heavily degraded source quality. Faded thermal receipts, scans under 200 DPI, or compressed images with artifacts. The AI compensates using context better than traditional OCR, but poor source quality is the biggest accuracy bottleneck regardless of document type.

Highly document-specific fields with no shared semantic anchor. If one document contains "Container Seal Number" and nothing else in the batch references something similar, the AI has no cross-document signal to anchor on. Type-specific fields extract better when batched with similar documents.

Unlabeled numeric values without nearby field names. A dollar figure sitting alone in a paragraph without "Total" or "Amount" nearby gives the AI no label to match against. Label-value pairs extract reliably; isolated numbers in narrative text may not.

Frequently Asked Questions

Can I batch-extract data from PDFs, scanned images, and phone photos all in the same upload — or do I need to separate them by format?

All formats go into the same batch. Upload PDFs, JPGs, PNGs, WebP, and AVIF files together. The visual language model processes every file through one pipeline — it reads the pixels, understands the content, and extracts the columns you named. A Document Date found in a native PDF is found the same way in a scanned JPG or a mobile photo: by semantic recognition, not file format detection. No separate batches per format, no manual format pre-sorting.

What happens if a field I defined — like Tax Amount — doesn't exist on some documents in the batch?

The AI leaves that cell empty for that document rather than fabricating a value or halting the batch. A purchase order without a tax line will show a blank cell in the Tax Amount column, while invoices in the same batch with tax will have the value filled in. No error stops the batch, and no fabricated value silently enters your spreadsheet. This is especially useful when your batch contains different document types or vendor formats that carry overlapping but not identical fields — each document is processed independently within the same column framework.

Can I extract line items in batch mode when each document has a different number of rows — from a single item on one invoice to 50 rows on a purchase order?

Yes. The AI reads each document's structure independently and generates as many output rows as that document contains. A single-line invoice produces one row with the header fields (Document Date, Vendor, Total Amount) and one line item row (Quantity, Unit Price, Line Total). A 50-line purchase order produces 50 rows — all under the same column headers. You don't pre-define row counts and you don't get a broken table from variable-length content. Computed Columns like "Line Total (Qty × Unit Price)" also work across variable-length batches.

Can I add calculated columns to a batch extraction — for example, Line Total computed as Qty × Unit Price during the batch run?

Yes. Computed Columns let you define calculations that the AI performs during extraction rather than requiring a separate Excel formula step. Type "Line Total (Qty × Unit Price)" as a column name, and the AI multiplies the values it finds for Quantity and Unit Price on each document in the batch, outputting the result directly. This works across every document regardless of how many line items each one has. For logged-in users, Rule Format supports more complex multi-step calculations including cross-row aggregation and conditional logic.

What happens when one document in the batch is poor quality — does it affect the other documents?

No. Each document is processed independently with its own status tracked in the batch view. If one file in a batch of 30 is a faded scan, that document's extraction accuracy will be lower — but the other 29 files are unaffected. The poor-quality document still appears in the output with whatever values the AI could extract, and its row is clearly identifiable for review. For best results, scan at 300 DPI or higher, and use the per-document status in the batch view to identify which files need attention without re-running the entire batch.

📮 contact email: [email protected]