Read Names, Numbers, and Readings from Photos of Physical Documents
Photographing a business card or parking ticket takes a second. Typing its data into a spreadsheet takes 2-3 minutes per item — this reads it in 5.
5-10s per photo · Up to 99% accuracy on printed text
What You Can Pull from Photos of Physical Documents
Business cards, shipping labels, utility meters, prescription labels — each one is a different physical object with different lighting, angles, and text quality. With Custom Column Extraction, you name the fields you need. The AI locates them on any photo by understanding what each value means, not where it appears in the image.
These are the column names you type for each scenario. The AI finds matching values regardless of the physical medium — one spreadsheet across all photo types.
Photos, Not Screenshots — Different Medium, Same Extraction Logic
A business card is glossy cardstock. A utility meter is a glass LCD panel. A prescription label is curved plastic. Each one reflects light differently, sits at a different angle, and carries text of a different size and font. Traditional tools need a separate configuration per medium — semantic extraction reads them all from one set of column names.
Why Photos Break Traditional Data Entry
Every physical object is a different reading challenge. Glossy business cards produce glare. Utility meter displays have low contrast in daylight. Prescription labels wrap around curved bottles. Traditional OCR expects flat, evenly-lit documents — physical objects never cooperate.
Manual transcription takes 2-3 minutes per item. A field worker photographs a meter, opens a spreadsheet, reads the number off the photo, types it in. Whiteboard notes are transcribed by hand. The gap between taking a photo and having data in a spreadsheet is filled by manual entry with every item.
Template tools need one config per medium. A zonal OCR tool needs extraction zones drawn for a business card, then entirely different zones for a meter photo. The physical variety makes template maintenance impractical for anyone handling more than one document type.
How Semantic Extraction Reads Any Physical Medium
You name the columns — AI finds values by meaning. Type "Name, Phone, Company" for a batch of business cards, then add "Reading Value, Meter ID" for meter photos. The visual model locates each field by understanding what it represents — a name is always a name, whether printed on glossy cardstock or written on a whiteboard.
One column set handles different physical media in the same batch. Upload business cards, shipping labels, and prescription labels together into one batch. Define columns once. The AI handles the glare on the cardstock, the curved surface of the bottle, and the small font on the label independently — every photo becomes a row in the same spreadsheet, even though no two photos share a physical format.
Visual context resolves lighting and angle distortions. A meter reading photographed from below at dusk — the visual model interprets the digits through the glare and low contrast, because it understands the concept of "a meter reading" — a numeric value framed by a label or unit indicator — rather than trying to cleanly OCR individual characters.
From Mixed Photo Types to One Spreadsheet in Three Steps
Photograph or Upload Physical Documents
You're at a trade show with a stack of business cards, at home with a utility meter, or at the pharmacy with a prescription bottle. Snap photos with your phone — JPG or PNG. Drag them into one batch: business cards, a shipping label from a package, and last month's meter reading. No need to sort by type — they all go in together.
Name the Fields You Need — Once
Type Name, Phone, Company, Tracking Number, Reading Value, Medication, Dosage. One set of columns covers every photo in the batch. The AI reads each image independently — it locates a phone number on a business card by understanding phone number patterns, reads the tracking number from a shipping label barcode area, and identifies the meter reading from its unit indicator. You don't tell it which photo is which.
Download a Single Clean Spreadsheet
Processing takes 5-10 seconds per photo. The output is one XLSX file: each row is one photo, each column is one field you named. Business cards in rows 1-5, the shipping label in row 6, the meter reading in row 7 — all in the same table. Roughly 18x faster than opening each photo and typing the data by hand (~2-3 min manual per item vs ~5s here).
When It Works Best — and When to Be Cautious
Understanding these boundaries helps you get consistent results.
When It Works Best
Well-lit, direct photos of flat objects. Business cards on a table, shipping labels on a desk, ID cards held flat — up to 99% accuracy.
Digital display readings with clear numerals. Utility meters, glucose monitors, and blood pressure devices use high-contrast LCDs. The model reads these consistently across brands.
Mixed-type batch processing. Business cards, labels, and meters in one batch with one column set — fields absent on a given photo stay empty. No per-type sorting needed.
When to Be Cautious
Extreme angles, reflections, or low light. A utility meter shot from ground level at dusk or a glossy card with overhead reflections will reduce accuracy. Photograph objects straight-on and in even light.
Heavy cursive handwriting on whiteboards or recipes. Printed handwriting extracts well; loose cursive or faded marker is less reliable. Treat handwritten output as a draft and spot-check.
Visible data only — no clinical interpretation. A reading of 180 mg/dL on a glucose meter is extracted accurately, but the AI does not flag health concerns. Clinical interpretation remains your responsibility.
Frequently Asked Questions
Can I extract contact details from a business card photo and a tracking number from a shipping label in the same batch?
Yes. Define one set of column names — Name, Phone, Company, Tracking Number, Carrier — and the AI finds matching values on each photo independently. A business card has Name and Phone but no Tracking Number; the shipping label has Tracking Number and Carrier but no Phone. Fields absent on a particular photo stay empty. The output is a single spreadsheet where every photo type contributes its relevant data.
How accurate is the extraction for meter readings taken from a phone photo at an angle?
For a straight-on, well-lit photo of a digital meter display, accuracy on the numeric reading reaches up to 99%. Angled shots or glare reduce that — the visual model still outperforms traditional OCR because it understands the meter reading as a semantic concept (a numeric value next to a unit label like 'kWh' or 'm³'), rather than trying to cleanly OCR distorted characters. For critical billing readings, take the photo as squarely as possible. The model reads it in 5 seconds regardless.
Can the AI extract handwritten content from a whiteboard photo or a handwritten recipe card?
The visual model handles handwriting, but legibility is key. Clear, printed-style handwriting on a well-lit whiteboard or recipe card extracts reliably. As one Reddit user described in a discussion about OCR tools, the manual fallback is "retaking the photo, opening the file, and typing out what I scribbled" — the shortcut saves time even when accuracy isn't perfect. Loose cursive or faded marker reduces accuracy. For handwritten content, treat the output as a starting point and spot-check.
What specific fields can I pull from a prescription label photo?
You can define any column you need — common choices include Medication Name, Dosage (e.g., "500 mg"), Frequency ("Take 1 tablet daily"), Prescriber Name, Date Filled, Pharmacy, and Refills Remaining. The AI reads both the drug name and the associated instruction text, routing each value to the correct column by understanding the label context. The tool extracts whatever is printed on the label; it does not interpret clinical meaning or verify drug interactions — that remains the pharmacist's domain.
Can I extract readings from both a glucose meter and a blood pressure monitor in the same batch to track health data over time?
Yes. Upload photos of your glucose meter display and blood pressure monitor screen into one batch. Define columns like Date, Systolic, Diastolic, Pulse, Blood Sugar, and Notes. The AI fills in Systolic, Diastolic, and Pulse from the blood pressure photo and Blood Sugar from the glucose meter — unrelated fields stay empty for each row. A "Notes" column can capture labels like "before breakfast" or "evening reading" if those appear on the device screen. Export as XLSX and you have a unified health log, roughly 18x faster than typing each reading manually.
Deep dives into photo document extraction: Field-by-field accuracy analysis for shipping label extraction including tracking numbers, handwritten annotations, and manifest tables · AI meter reading technology compared with manual reading, AMR, and smart meters for utility scenarios · Prescription label extraction accuracy analysis for pharmacy teams evaluating AI for medication safety
If you need the entire document turned into a full structured table rather than just specific fields, screenshot-to-excel conversion handles a different type of extraction need — for screen captures rather than photos of physical objects.