Link Coffee Lot Invoices
to Producer Plots for Traceability
Coffee supply chain traceability doesn't break where most compliance articles say it does. It breaks at the moment someone tries to connect the lot number on a shipment invoice to the producer plot list on the supplier's traceability schedule, and the two documents use different identifiers, different spellings, or don't reference each other at all. Under the EU Deforestation Regulation (Regulation (EU) 2023/1115), which applies to large and medium coffee operators from 30 December 2026 and to micro and small enterprises from 30 June 2027, that link is the backbone of a due diligence statement, and it is a document problem long before it becomes a legal one.

Key Takeaways
- Coffee traceability breaks on the document link, not on the deforestation data most guides blame.
- The invoice says COFFEE-10 and the schedule says COF-10, so a person sees one lot while a string comparison sees two.
- ImageToTable.ai's Custom Column Extraction runs the invoice, certificate, and schedule through one column set, so a mismatch becomes a row you sort instead of a pile you hunt through.
What a Traceable Coffee Lot Actually Requires

Every bag of green coffee carries a statutory identity mark. Under the Rules on Statistics of the International Coffee Organization (ICO), each export parcel gets a three-part identification mark printed on every bag and on its certificate of origin: the producing country code, the exporter or grower code, and the parcel serial number. An Ethiopian lot marked 010/0123/0047 reads as Ethiopia, exporter 0123, parcel 0047.
That mark is the parcel's passport. The European Coffee Federation's standard contract for green coffee shipments, the ESCC, lists the supporting documents that travel with it in Article 20: the commercial invoice, the bill of lading, a certificate of weight, a certificate of origin, and phytosanitary and fumigation certificates. From the importer's side, one lot is therefore not one document. It is a small pile of documents that are supposed to point at the same parcel, the same origin, and the same quantity.
EUDR adds one thing the old paperwork never asked for: the link down to the plot. Under Regulation (EU) 2023/1115, the information an operator must collect includes the geolocation of every plot of land where the coffee was produced, with latitude and longitude at six decimal places, a point for plots up to 4 hectares and a polygon for larger ones, plus the harvest period. The European Commission's EUDR guidance is the reference point for how those geolocation details are applied and which records the due diligence statement must cover. The traceability schedule a supplier or cooperative sends with the lot is where that plot list lives. The invoice names the lot. The schedule names the plots. Traceability is the act of proving they are the same lot.
A traceable coffee lot is a bundle of documents, and lot traceability is the link between the lot number they all claim to describe.
The Documents Behind One Lot, and Who Handles Them
A single container of green coffee routinely combines coffee from hundreds of smallholder farms through a cooperative or mill, and the paperwork arrives from different hands at different times.
- The commercial invoice comes from the exporter and names the lot, the buyer, the quantity, the HS code (0901 for green and roasted coffee), and usually the country of production. It is the document the importer's own purchase order has to be reconciled against.
- The ICO certificate of origin is issued by the producing country's authorized body and carries the same three-part parcel mark plus quantity and destination. It is the "official" statement of origin, which is why auditors reach for it first.
- The traceability schedule or plot list links the lot to producer plots. It may be a table from the cooperative listing plot IDs, farmer names or codes, areas, geolocation references, and harvest windows, or a GeoJSON-style reference to geometry files held by the cooperative.
- Certificates (organic, Rainforest Alliance, Fairtrade, 4C) add chain-of-custody claims. The contract with the buyer usually requires the certificate number to appear on the invoice, so the certificate and the invoice must agree as well.
On the buying side, a compliance or sourcing team reconciles these documents before a shipment moves. On the producing side, an exporter assembles them from cooperative records that may still live in paper ledgers, spreadsheets, or a farm data platform. In coffee, a large share of the world's supply is grown by smallholders on plots under four hectares, so the plot-level data often exists but is scattered, in different formats, and keyed by different identifiers. For field-data and plot-type records, the extraction pattern is the same one used for agricultural and environmental survey forms, which we cover in our guide to extracting field data from ag and environmental forms.
The invoice, the certificate, and the schedule each answer a different question about the same lot, and they rarely answer it in the same words.
Where the Invoice-to-Plot Link Breaks in Practice

Practitioners describe the failure modes far more precisely than any marketing page does. A new coffee importer on r/coffee_roasters summed up the structural problem: "how unstructured relationships can be from origin all the way to the roaster, even with all the technology available today." On r/Coffee, another participant put the cost on the table: "Traceability is expensive at every level of the supply chain." Those are the two real forces at work: unstructured records and expensive linking.
Concretely, the link breaks in a handful of recurring patterns:
- The invoice lot number and the schedule lot number differ. The invoice says COFFEE-10, the traceability schedule says COF-10 or drops the leading characters. A human can see they are the same lot; a string comparison cannot, and a customs or audit review will treat them as two different lots.
- The port of loading is mistaken for the country of production. EUDR's origin tests the plot where the coffee was grown, not the port it sailed from. A schedule that lists "Mombasa" as the origin because that is where the washing station shipped it fails the check a compliance team actually runs.
- Plots on the schedule have no geolocation reference. The schedule names 120 plots but only 80 carry coordinates or a geometry file reference. Every listed plot needs a geographic evidence reference; the gaps only surface when someone lines up the plot list against the evidence.
- Documents arrive as a mixed pile of PDFs, scans, and photos. Invoices from the export house, certificates scanned at a coffee shop in origin, the schedule as an emailed spreadsheet. Each arrives in a different format, which is precisely the case where template-based tools stop working.
- Yield and weights disagree across documents. The invoice says 19,200 kg net, the certificate says 19,195 kg, and the packing list says something else again. None of these are an EUDR problem on their own, but they are the discrepancies that make an authority look twice.
Every one of these failures is a document-matching failure before it is a compliance failure. The lot exists; the paperwork disagrees about which lot it is.
The good news: none of these mismatches require judgment about deforestation or legality. They require taking the identifiers and references out of the documents and putting them next to each other, which is exactly what a spreadsheet is for. The hard part is getting them out of documents that use a different layout every time. The same cross-document reconciliation mechanics that procurement teams use for supplier invoice and purchase order matching apply here, and we explain that pattern in depth in supplier invoice to PO matching in manufacturing.
Pull Every Lot Field into One Spreadsheet, Then Compare

The workable pattern is to define the columns you need once, process every document in the lot's file pile with the same column set, and export everything into a single table. Because every document type is extracted against the same columns, the lot number from the invoice ends up in the same column as the lot number from the schedule, and the comparison becomes reading rows instead of hunting through PDFs.
The mechanism is Custom Column Extraction: you type the column names you want, and the AI locates each value anywhere on the page by understanding what it means, not where it sits. No templates, no training, and no need to pre-configure a layout for each supplier. A column set for one coffee lot might look like this:
Sample column set for a traceability reconciliation table
- Source Document ID (extracted, to see which document each row came from
- Document Type (inferred column with options: Invoice / Certificate of Origin / Traceability Schedule / Certificate
- Lot Number (extracted, from the invoice or the schedule
- Country of Production (extracted
- Plot ID (extracted for schedule rows; blank for invoice rows)
- Geolocation Reference Present (inferred Yes/No on whether the plot row carries coordinates or a geometry reference)
- Plot Count per Lot (computed as a count of schedule rows grouped under the same lot)
- Net Weight (extracted where present)
Processing the invoice, the certificate, and the schedule through the same column set produces a table where every row is one document, and every lot number sits in one column. Sorting by lot number lines up the invoice row, the certificate row, and the schedule rows for the same shipment, and the mismatches become visible: a lot number that only appears once, a schedule whose plot rows mostly read "No" under Geolocation Reference Present, a certificate that names a different origin than the invoice.
Because the extracted values stay traceable to their source location, the review step stays honest. Review Mode with bbox highlighting shows exactly which part of the original document each extracted value came from, so when two documents disagree, you can see the printed values behind both claims instead of trusting the extraction. That matters for a compliance file, where "the AI said so" is not an answer an auditor accepts.
Try the extraction on one of your own documents in the demo below, with the same column names you would use for your traceability table.
Files are processed securely and not stored.
Collecting the File Pile from Suppliers Without Chasing Inbox Threads
Assembling the pile is half the work in traceability, because the invoice arrives from the export house, the certificate from a broker, and the schedule from the cooperative, each in its own email. Two product mechanisms make the gathering step repeatable instead of a per-shipment scramble.
Collection Link generates a link you send to the supplier or cooperative; they open it, enter a short code, and upload their documents directly into your processing queue with no account of their own. Email Inbox gives the account a dedicated address suppliers can forward invoices and certificates to, with an optional sender whitelist so only approved partners can drop files into the queue. Attachments land automatically and a bound template starts extraction the moment mail arrives. Both paths turn "please email it to us" into a structured intake that lands in the same batch as everything else.
The acquisition step and the extraction step are the same workflow: documents stream in through a collection link or a mailbox, and every inbound file is processed against the same traceability column set. For a field-level breakdown of pulling identifiers out of purchase documents, see our complete guide to purchase order data extraction, and for the accounting side of feeding extracted document values into production cost tracking, our guide on supplier invoice to production cost tracking covers the invoiced-quantity trail that matters when the same lots move through roasting and blending.
What Stays Outside the Spreadsheet
It is just as important to say what this workflow does not do, because a compliance file built on the wrong assumptions is worse than a slow one.
- We do not determine EUDR compliance. Comparing lot numbers and origins proves the documents describe the same lot. Whether that lot is deforestation-free and legal is a conclusion your compliance team or adviser makes from the plot evidence, and it cannot be automated away.
- We do not validate geometry or coordinates. A geolocation reference being present is extracted; validating that a polygon is closed, that coordinates land on the right plot, or that they survive a satellite forest-cover check is GIS work best done in tools made for it, plus the verification steps a competent authority expects.
- We do not file the due diligence statement. Submitting the DDS in the EU Information System (TRACES NT) is the operator's legal act and has to go through the official system. The spreadsheet is the working file behind it, not the filing itself.
- Computed columns reason within one row or one document. The computation you can ask for during extraction is the kind that operates on extracted values in the same document, such as counting the plot rows under a lot. Deciding that two separate documents disagree is judgment you keep in the spreadsheet review, which is exactly where we put it above.
This division of labor is not a limitation of the tool. It is the correct distribution: machines are reliable at pulling identifiers out of unfamiliar layouts, and humans (or their compliance advisers) are responsible for the conclusion those identifiers support. Building documentation systems that respect that boundary is the difference between a traceability file that survives audit and one that only looks complete.
A Field-by-Field Checklist for the Next Shipment
Once the columns are set, tracing coffee lots to producer plots works the same way on every shipment, and a repeatable reconciliation workflow comes down to a checklist you can run in an afternoon once the documents are extracted side by side.
- Confirm the lot number matches across the invoice and the schedule. For each shipment, the lot number column should contain the same value on the invoice, the certificate, and the schedule rows. Any value appearing only once is a flag to resolve with the supplier.
- Confirm the country of production on the invoice matches the certificate. If the invoice names one country and the certificate another, resolve before the file goes anywhere near a statement.
- Confirm the port on the bill of lading is not used as the country of production. The origin field must describe where the coffee was grown.
- Check every plot on the schedule for a geolocation reference. Sort the schedule rows by the Geolocation Reference Present column, and chase the rows that read "No" with the cooperative before shipment, while there is still time to fix them.
- Reconcile quantities. Compare net weight on the invoice, the weight certificate, and the packing list. Discrepancies that tiny (19,200 vs 19,195 kg) are the ones that annoy customs officers.
- Archive the working table with the evidence file. EUDR requires operators to keep due diligence information for five years. The extracted spreadsheet, with every value live-linked to its source document, is a far easier thing to produce at an audit than a folder of forwarded emails.
Run this checklist at contract time and at each shipment, not the week a statement is due.
FAQ
Can I compare the invoice to the traceability schedule automatically?
You can bring both documents into one spreadsheet through the same column set, which lines up identical fields like lot number and country of production so mismatches are visible at a glance. Automatically deciding that two documents refer to the same lot is kept as judgment someone reviews, because that decision is what a compliance conclusion hangs on. The extraction does the reliable part (reading the values); the review does the accountable part (deciding what they mean together).
Does this work when every supplier sends a different format?
Yes. Custom Column Extraction reads by meaning rather than by fixed layout, so an invoice from a trading house, a scanned certificate, and a spreadsheet-style schedule all go through the same columns without per-supplier templates. That is the main difference from template-based OCR tools, which need a template per layout and fail when a supplier updates theirs.
Can extracted data be used in the EU Information System (TRACES)?
The extracted table gives you the structured values that feed a due diligence statement, but filing the DDS must happen through the official EU system. The spreadsheet is the working file; the submission is a separate legal act by the operator. Our role stops at producing clean, verifiable data from the documents.
Does an EUDR due diligence statement apply to me as a roaster or trader?
The obligation depends on your role in the chain, not on your job title. Whoever first places the coffee on the EU market or exports it (typically the importer for green coffee) files the due diligence statement and carries the full information requirements. Downstream operators and traders keep lighter obligations, but they still record and retain their suppliers' statement reference numbers and the related data. For load-specific and role-specific deadlines, check the current guidance on the European Commission's deforestation regulation page, since the application timeline has been adjusted more than once.
Is a certification (Rainforest Alliance, Fairtrade, 4C) enough for traceability?
Certification is complementary evidence. Schemes are increasingly building plot-level data into their systems, and a certificate a buyer asks for on the invoice is part of the file. But certification alone does not replace the operator's due diligence, and the linked lot and plot data still need to be collected and checked. Treat certificates as supporting records in the same spreadsheet, not as a substitute for the reconciliation.
The Bottom Line
The invoice-to-plot link is a document problem, and it is the part of EUDR that no regulatory test can automate away: no authority can file your due diligence statement, and no geolocation tool can decide that a lot number on an invoice is the same lot a cooperative listed in its schedule. Someone has to put those documents next to each other and look. When the extraction step is replaced with a column-based pull that reads any layout, the reconciliation that used to cost hours of copy-paste turns into sorting one table and reading a checklist, and the evidence file behind the statement becomes something you can produce on request instead of reconstruct in a panic.
Test the extraction step on your own coffee documents below, and see how fast an invoice or a traceability schedule turns into rows you can actually compare.