Reading a Mill Test Report from a Bad Scan
What Actually Survives
The part of a mill test report that fails first on a poor scan is not the field label. It is the geometry of the chemistry table. A certificate can read cleanly line by line and still come apart the moment one element column drifts, because on an EN 10204 test report a percentage has no meaning until it is bound to the element header above it. The scan is also the one variable nobody in this chain controls. A mill issues a clean digital certificate, a distributor prints and faxes it, a service center photographs the fax into the receiving folder, and by the time it reaches the goods-in desk it is three generations from the original.

Key Takeaways
- A scan can read every word on a mill test report and still produce a row that links to the wrong material, because a value means nothing until it is bound to its label.
- When a poor scan merges two column rules, the percentages stay readable but inherit the wrong element labels, which is worse than a blank because the record looks complete.
- The useful question is which of the eight field groups came back clean, which is why ImageToTable.ai lets you name the columns and finds each value by meaning instead of by its position on the page.
What You Are Actually Pulling Off a Mill Test Report

A mill test report is not one document, and the EN 10204 certificate type printed on it tells you which fields are real measured values and which are only assurances. The standard defines four types, and the type changes what is on the page. Type 2.1 is a declaration of compliance with no test results at all. Type 2.2 is a test report built on non-specific inspection, which means the numbers may come from a different heat than the one delivered. Type 3.1 carries specific test results for the actual delivered heat, validated by the manufacturer's inspection representative who is independent of production. Type 3.2 adds an independent third-party witness on top of 3.1. Most structural and pressure-equipment steel ships as 3.1, so that is the certificate most receiving teams are reading day to day.
The field layout of that certificate is not arbitrary either. EN 10168 defines the coded sections a steel inspection document uses: Group A for commercial data, B for product description, C for inspection (chemical and mechanical), D for other tests, and Z for validation. Yield strength sits at section C11, tensile strength at C12, elongation at C13, notched-bar impact across C40 to C43, and the chemical analysis from C71 upward, where each element is named individually rather than assigned to a fixed slot. The fields you actually extract all live inside those sections: heat or cast number, grade, dimensions, heat treatment condition, element percentages from carbon to molybdenum, yield and tensile strength, elongation, impact energy, the certificate type itself, and the inspector's signature. The full list of certificate types and what each one requires is set out in the EN 10204 reference.
That coded structure is what makes a mill certificate a strong candidate for automated extraction, and it is also exactly what a bad scan destroys first. The goal behind extracting it rarely changes. In a 2015 r/manufacturing thread on MTR database software, a team described wanting to "import and store scanned MTRs, index them by item/shape/fitting/country, and tag them with where we have used them." That is the same traceability record receiving teams still need today, and it only holds if the fields come off the page intact.
Where a Poor Scan Actually Costs You

Scan quality does not degrade a mill test report evenly. It attacks the fields that carry the most traceability weight.
Resolution. The digitization guidelines used for professional archival work put text capture at 300 ppi for three-star quality and 400 ppi for four-star, according to the FADGI Technical Guidelines, third edition. A fax page arrives at roughly 200 dpi in monochrome, and a phone photo arrives at whatever the operator managed. Between those numbers sit the hairline rules that separate one element column from the next, and the small-font superscripts that record a ladle analysis versus a product analysis. When the rules dissolve, the column structure dissolves with them.
Skew and rotation. A 2024 study of multimodal document models found that key-value extraction held steady to about 25 degrees of rotation for one model and 35 for another, then collapsed into hallucinated values rather than clean errors, as reported in research on skewed document extraction. Those limits look forgiving on their own. On a chemistry table with a dozen columns, the binding constraint is not page rotation but vertical-rule alignment, where a small tilt plus scanner speckle merges adjacent rules long before the page itself looks crooked. FADGI treats four-star capture as a tolerance of plus or minus 1 degree with no software deskew permitted, which is an alignment a fax or a handheld photo almost never reaches.
Contrast and background noise. A photocopy of a fax inherits both problems at once: grey paper, a speckled background, and letterforms whose strokes have swollen or thinned. This is where heat numbers with visually similar characters become the most common misread, because 0 against O, 1 against I, and 8 against B differ by a single pixel at low resolution. A wrong digit in a traceability key is not a typo anyone can see in the output. It is a row that links to the wrong material.
Stamps, annotations, and merged cells. Rubber stamps over a heat number, handwritten corrections over a printed value, and chemistry tables whose header spans merged cells all break the pairing between a value and its label. A linear read reports the words and loses the binding. The number is on the page, but you no longer know which element it belongs to, which is worse than a blank because it looks complete.
Multi-heat tables. A single certificate often covers several heats, laid out as one row per heat. When the row structure is lost to compression or a fold line, values from different heats blend into one record, and the resulting chemistry looks plausible while belonging to no real heat at all.
The chemistry table is the first casualty of a bad scan, and a dropped element column does not produce one missing number. Every value to its right inherits the wrong label.
What Survives, Field by Field

Heat numbers and certificate types survive a poor scan better than chemistry-table labels, and signatures survive worst. The difference is not luck. It follows from how each field is printed and how much redundancy the document gives it.
| Field | Behavior on a poor scan | What to do |
|---|---|---|
| Certificate type (2.1 / 2.2 / 3.1 / 3.2) | Printed large in the header and often repeated | Reads on the first pass |
| Heat or cast number | Short and high contrast, but 0/O, 1/I, 8/B confusable | Verify against the material marking |
| Grade and standard | Alphanumeric, repeated in header and product section | Check the edition or revision |
| Chemistry percentages | The numerals survive, the element labels are bound to the header | Confirm column alignment |
| Yield, tensile, elongation | Usually printed beside the property name | Reliable on printed certificates |
| Impact energy and hardness | Often inline with a test temperature | Verify units and temperature |
| Heat treatment and NDT notes | Prose and abbreviations, sometimes stamped | Expect gaps on faint copies |
| Inspector signature and stamp | Ink over print, faint by nature | Human check, never trust blind |
The pattern is consistent across mills: fields with redundancy and large type survive, and fields that depend on physical position or ink quality do not. That is why the useful question is not "did the tool read the certificate" but "which of these eight field groups came back clean, and which ones do I need to look at."
The Fix: You Name the Fields, and the Tool Finds Them
The reason a field-named extraction survives a bad certificate is that it does not need the table rules to be intact, only the element name and its value to be readable somewhere on the page. ImageToTable.ai works this way through Custom Column Extraction: you type the column names you want, and a vision model reads the certificate to locate each value by the meaning of the field rather than by coordinates. On a mill certificate, a column called "Heat Number" is found whether the mill prints it as "Heat No.", "Cast No.", or a foreign-language equivalent, and a column called "Yield Strength (MPa)" is found whether it sits on its own row or inside a test block.
Two things make the output trustworthy enough to enter a quality record. Batch processing lets you drop a week of certificates and get one spreadsheet with a row per certificate or per heat. The review layer then maps every extracted cell back to its position on the original scan: Bbox verification highlights exactly where a value came from when you hover it, so a yield strength that looks off can be checked against its source in one click instead of by eye across the whole page.
Model Tier is the other half of the answer for bad scans. ImageToTable.ai accounts run at a Standard, Advanced, or Premium processing tier, and higher tiers use a stronger underlying vision model for dense tables, faint faxes, and handwriting. A clean digital certificate from a major mill is a Standard job. A photographed, stamped, multi-heat certificate from a regional mill is the case for Advanced or Premium. Whatever tier is active when a batch is submitted is what that batch is billed against, so you can send the difficult pile at a higher tier and keep the easy pile economical. For a denser chemistry table where you also want a computed check, a Computed Column can output a derived value from the fields already extracted, such as a ratio or a difference, without retyping anything in Excel.
Files are processed securely and not stored.
What This Will Not Do
Extraction gets the numbers off the page, and it stops there. Everything that requires judgment about whether those numbers are acceptable stays with your QA team. That line is worth drawing clearly, because a tool that quietly crosses it is more dangerous than one that does not.
- It does not validate chemistry or mechanical results against grade limits. Checking a carbon percentage against ASME Section II, EN 10025, or your project specification is a rule your quality system applies to extracted data. Extraction supplies the values; it does not decide whether the values pass.
- It cannot recover a genuinely illegible heat number. A digit fused by a stamp or lost to blur should be flagged for review, never guessed. A wrong heat number is more expensive than a blank one, because it points traceability at the wrong batch and no downstream check catches it.
- It does not treat a multi-heat certificate as several records unless the row structure is readable. Verify one row per heat before import when a certificate covers multiple heats.
- It does not prove that the steel on the truck matches the certificate. Comparing certificates to each other only proves the paperwork agrees with itself. Receiving inspection closes part of that gap, and on critical service, positive material identification (PMI) closes the rest.
- It does not authenticate a signature or confirm that a type 3.2 third-party counter-signature is present. That is certificate-of-conformance review, and it needs a person.
A Goods-In Workflow That Holds Up
A workflow that respects both the traceability requirement and the reality of bad scans looks like this, and every step maps to a specific setting rather than a general promise.
Define your columns from the EN 10168 sections
Name the fields you actually track: Heat Number, Grade, Certificate Type, C, Mn, Si, P, S, Yield Strength (MPa), Tensile Strength (MPa), Elongation (%), Impact Energy (J), Heat Treatment. The column names you enter become the headers of the output sheet, so decide them once for all mills instead of rebuilding a format per supplier.
Split the pile by scan quality
Sort the week's certificates into clean digital PDFs and degraded paper (faxed, photocopied, photographed). Run the clean stack at the Standard tier and the degraded stack at Advanced or Premium. Tier is set per account and billed per submitted batch, so the split costs no more than the documents need.
Upload each stack as one batch
Batch processing takes many certificates at once and merges them into a single spreadsheet, one row per certificate or per heat. Upload the whole stack with the shared column set instead of processing certificates one at a time.
Review the fields the scan threatens, using Bbox
Treat heat numbers, chemistry column alignment, and anything under a stamp as the review set. Hover a suspect cell to see exactly where the value came from on the original, correct it if needed, and revert to the AI value if the original was right after all. The fields with redundancy, like grade and certificate type, usually need no more than a glance.
Export and import into the system that needs it
Export to Excel, CSV, or JSON and bring the file into your ERP or quality system through its normal import path: Epicor Kinetic, Infor, SAP Business One, Microsoft Dynamics 365, and SYSPRO all accept file-based imports with a column mapping. The extraction produces the data layer; the ERP holds the record.
For a worked example of the receiving-inspection side of this same data problem, our guide to extracting quality inspection report data covers dimensional measurements and disposition fields. If your degraded documents are forms rather than certificates, the low-quality scanned forms workflow handles the same mechanics from a different starting point, and the inspection report extraction hub maps the full set of report types. If you are still choosing tools for a plant that handles several document types, the manufacturing extraction tool comparison puts MTRs in the context of purchase orders, packing slips, and inspection forms, and the field and industrial tool comparison covers the dock-and-yard side.
FAQ
Can it read a faxed or photocopied mill test report?
Often, yes, and the honest answer depends on how many generations of copying the page has survived. A first-generation fax at around 200 dpi usually keeps the header fields and the printed chemistry numerals readable, while a photocopy of a fax loses the hairline table rules and the element labels with them. Use a higher Model Tier for this pile, and expect to review the chemistry alignment even when the values come back. A scan that is clean when it left the mill does not need any of this.
Does it understand EN 10204 certificate types like 3.1 and 2.2?
It extracts the certificate type as a field, so you can carry it into your records and filter on it. The type is usually printed large in the header and reads reliably even on poor copies. What it does not do is enforce your acceptance rule, for example refusing a 2.2 where your purchase order specifies 3.1. That is a downstream check in your quality system, built on the extracted value.
Will it validate the chemistry against the material grade?
No, and this is deliberate. Validation means comparing extracted values against the allowable ranges for a grade in ASME Section II, EN 10025, or your own specification library, and routing out-of-range results. That is a rule your quality team owns and maintains because the limits change with grade, thickness, and standard edition. Extraction gives you the values in typed columns so that check becomes a formula or a rule, not a manual read of every certificate.
What file formats does it accept, and how many at once?
PDF, JPG, PNG, WebP, and AVIF, including password-protected PDFs. Batch processing takes many files in one submission and merges them into a single spreadsheet, so a week of certificates lands as one dataset with a row per certificate or per heat. Photos of paper certificates work, provided the image is reasonably sharp and evenly lit.
Can it handle a certificate that covers multiple heats?
It can return multiple rows from one certificate when the heat rows are readable, which is what you want for a rebar or heavy-plate delivery covering several heats. When the row structure is damaged by compression or a fold, the rows can blend, so a multi-heat certificate belongs in your review set, and you should verify one row per heat before the data reaches inventory.
Is handwriting on a certificate a problem?
Handwritten values and annotations are harder than printed ones and are exactly where the higher Model Tiers earn their cost. Clear block capitals and standard abbreviations extract well. Dense cursive, smudged corrections, or a value written over a printed figure should be treated as a review item, especially when the handwritten figure is a heat number or a measured value that carries compliance weight.
The useful mental model is a division of labor: extraction owns the page, and your quality system owns the judgment. A bad scan removes information, and no model can put back a digit that a stamp has fused or a column that a fax has erased. What a field-named extraction does is tell you precisely which fields came back clean and which ones sit on damaged ground, so review time goes to the heat number under the stamp instead of the grade that every certificate prints in the same place.
Test it on your worst certificates rather than your cleanest PDFs. Take a photographed, multi-heat certificate with a stamp over the heat number and see which columns come back and which ones flag for review. Drop a few in and check the alignment yourself.