Your Vendor List Has 200 Names
About 120 Real Suppliers
A procurement lead at a 140-person company was handed a vendor consolidation project and a starting point from finance: a list of more than 200 vendors paid in the past year. The list existed. A usable picture of spend did not. The same supplier appeared "4 different ways," as the poster put it, and on the reimbursement records the vendor field held employee names instead of company names. A commenter on the thread guessed that the 200 rows were probably about 120 real suppliers (r/procurement).
That gap between 200 and 120 is the whole problem. You cannot deduplicate a vendor list until the vendor name becomes a reliable key, and after a year of purchases on personal cards with no central tracking, it is not one yet.

Key Takeaways
- 200 vendor names usually resolve to about 120 real suppliers, and that gap is not a data-entry slip but the default outcome of a year of personal-card buying.
- "AMZN MKTP US" and "Amazon.com" are the same supplier, but no fuzzy match connects them, because the truncated descriptor shares almost none of the brand name's characters.
- There is no vendor master to clean, only receipts and statements that were never in one table, so before matching names, put every record into one sheet with the same columns (ImageToTable.ai reads by column name, not layout).
The Consolidation Deadline Meets a List Nobody Can Use

A vendor list from finance tells you who got paid. It rarely tells you who you are actually buying from. Vendor spend consolidation depends on that second question, and those sound like the same thing until you try to sum spend by supplier and find that no two rows for the same company agree on the name.
The data is complete enough. The problem is that the key is unreliable, and every consolidation question sits downstream of that key.
Ranking suppliers by total spend requires a stable name. Spotting that three teams pay for overlapping project management software requires a stable name. Checking whether a negotiated rate is actually being used requires a stable name. When the key is a name that was typed by a different person on a different day for each purchase, none of those questions can be answered, and the negotiation leverage you were supposed to create never materializes.
The cost of getting this wrong is not abstract. APQC's Open Standards Benchmarking puts the median organization at duplicate or erroneous payments equal to 1.5% of annual disbursements, with top performers still at 0.8% (APQC). APQC names poor-quality data in the master vendor file as one of the core causes (this is the same class of list you are trying to build from scratch). A company with $10M in spend is looking at six figures of leakage that better supplier visibility would have caught.
Where the Records Actually Live

In a company without a purchasing system, the vendor record is not in one place. It is scattered across four sources that each name the supplier differently.
Card statements
Whatever the merchant descriptor says. This is a machine-generated string, not a name a human chose.
Receipts and invoice PDFs
The actual document, often a photo. Often the only place the real trading name appears.
Expense reports
Submitted by the employee who paid. The payment is real, but the "vendor" field on the report is frequently the employee's own name.
The finance export
A spreadsheet assembled from the above. It covers what was reimbursed or entered, not everything that was bought.
So the consolidation task is not "clean the vendor master." There is no vendor master. The task is to build one from the raw records first, and only then start the name-matching work everyone assumes you began with.
Four Reasons One Supplier Appears Under Four Names

Vendor name variance has mechanical causes, and knowing them tells you which variants are safe to merge and which need a human decision.
Nobody owned the name. With decentralized buying, the person filing an expense chooses what to type. One employee writes "Acme Supply," the next writes "ACME SUPPLY LLC," a third writes "Acme." Exact-match deduplication, the kind Excel does with conditional formatting or a COUNTIF, returns almost nothing, because no two strings are identical.
The card descriptor is truncated and encoded. Card networks cap the merchant descriptor (the limit is commonly 22 characters), and processors prefix it with their own tag. A purchase from Amazon can land on a statement as AMZN MKTP US, a payment through Square as SQ *MERCHANT. The brand name is gone, replaced by a code. A receipt for the same purchase says "Amazon.com." No string-similarity algorithm will connect those two, because they share almost no characters.
Legal name, DBA, and remit-to are three different strings. A supplier's legal entity may be "Northwind Logistics Holdings LLC," its trade name may be "Northwind," and its remit-to address may belong to a factoring company or a payment processor. IOFM notes that the DBA field on a W-9 exists precisely so a buyer can match an invoice name that differs from the legal name (IOFM). Large suppliers also invoice from divisions, so the same parent company can appear as three separate vendors with the same tax ID.
Reimbursements record the payer, not the payee. When an employee pays and gets reimbursed, the transaction in the expense system is tied to the employee. The vendor may appear only on the receipt image, or nowhere at all if the receipt is missing.
Four mechanisms, four names. Add case differences, punctuation, and suffixes like Inc versus Incorporated, and the 200-to-120 ratio stops looking like an outlier and starts looking like the default outcome.
What Fuzzy Matching Can and Cannot Fix
Fuzzy matching is the standard next move, and it solves part of the problem. It scores how similar two strings are rather than requiring them to be identical. The usual measures are edit distance (Levenshtein, where the score is the number of character changes) and token-based similarity (Power Query's fuzzy merge uses the Jaccard similarity algorithm, per Microsoft's documentation). You set a similarity threshold, and anything above it is flagged as a possible match.
Fuzzy matching generates candidates. It does not make decisions, and the threshold you pick decides which error you prefer: misses or false merges.
The threshold is where it breaks down. Lower it enough to catch "ABC Supply Inc." and "ABC Supply LLC," and it starts merging unrelated names that happen to share tokens. An Excel University reader documented the failure precisely: at a 0.9 threshold the tool matched "Titan" to "Twitch" and "SAVE" to "Pave," and dropping to roughly 0.5 to catch the Inc./LLC variants produced far too many false positives to review (Excel University). You cannot tune your way out of that trade-off with names alone, which is why production deduplication weighs name against supporting attributes such as address, tax ID, bank details, or transaction pattern.
Two limits matter more here than the threshold. First, fuzzy matching fails outright on descriptor codes, because "AMZN MKTP US" and "Amazon.com" are not similar strings. Second, it cannot decide whether three divisions of one parent should be one vendor or three. That depends on whether you negotiate with them as one relationship or three, and that is a business decision, not a string comparison.
This is also where the work differs from catching a duplicate payment. Detecting that the same invoice was paid twice is a comparison against history, and we cover those failure modes in duplicate invoice detection. Building one clean vendor list from messy, decentralized records is a grouping problem across a whole population, and it has to be solved before any duplicate-payment check can be trusted, because a duplicate check that keys on an inconsistent name misses the very pairs it was built to find.
Both fuzzy matching and every other cleanup technique assume one thing: that all the records are already in a single table with a usable vendor column. After a year of personal-card purchasing, they are not. That is the step to fix first.
Step One: Get Every Record Into One Sheet With the Same Columns
Before any name can be normalized, every receipt, invoice, and statement line has to exist in one place under the same column headings. That is a data extraction job, and it is worth doing in a way that does not depend on the document's layout, because you are dealing with photos, scans, and PDFs from dozens of merchants, all formatted differently.
ImageToTable.ai uses Custom Column Extraction: you type the column names you want, and the AI locates the matching value anywhere on the page by understanding what the field means rather than where it sits. For this job the column set is small and stable:
Vendor Name(exactly as printed on the document or descriptor)Transaction Date (YYYY-MM-DD)AmountCard(which card the purchase was made on)Category (options: Software/Office/Travel/Meals/Other)
Two of those columns do more work than they look like they do. The date format instruction, which the tool calls a Format Requirement, makes every supplier's date land as the same string, so sorting and period filtering actually work on the first try. The Category column is an example of an Inferred Column, where the AI fills in a value that is not printed on the document by reading the receipt content and choosing from the options you supplied. That is how classification rides along with extraction instead of becoming a separate pass.
Because the tool is batch-first, you upload the whole folder and get one merged sheet rather than one file at a time. Card statements spanning several pages are handled by Multi-Page Merge, a template setting that groups pages belonging to the same statement into a single record so account-level information carries across the line items. If you also have year-end volumes to work through, the same batch-and-merge flow is covered in more detail in batch credit card statement processing.
Where you start depends on what the documents are. If most of the pile is receipt photos and email invoices, the receipt to Excel workflow is the fastest entry point. If it is a stack of monthly statements, start with the credit card statement extraction page instead. Either way, the output is the same shape of table, which is the point.
Files are processed securely and not stored.
One thing the output is not: a clean vendor list. Extraction gives you every record in one table with the vendor name as it appears in the source. That is the raw material for the judgment step. Verifying those names before you trust them is easy, since hovering or clicking any extracted cell highlights the exact spot on the original image it came from, and clicking a region on the image jumps back to the matching cell. This is the Review Mode with Bbox verification, and it is what lets you confirm that an odd string is genuinely what the document said rather than a misread.
Step Two: Build the Canonical Mapping by Hand, With the Data in Front of You
Vendor name normalization is a judgment task, and the honest answer is that you do it yourself, with the extracted table sorted so the decision order is obvious. Sort by total amount descending and work from the top.
For each supplier, decide what the canonical name is, then map every variant to it in a second column. A workable structure is three columns: the raw Vendor Name from extraction, a Vendor (canonical) column you fill, and an alias note for the variants you folded in. That preserves the original string for audit while giving you something stable to pivot on.
The top of the list resolves quickly. Eight design agencies become eight canonical names. Three overlapping project management tools become three, and now you can see the overlap and decide to cut one. A cluster of statement lines reading AMZN MKTP US, AMAZON.COM, and AMZN Prime all belong to Amazon, but AWS is infrastructure spend and usually belongs on its own line, so the grouping decision is not just "merge anything similar." That distinction is exactly the one fuzzy matching cannot make, and it is why you review the mapping instead of accepting an algorithm's output.
Use the supporting columns to confirm rather than guess. If two variants share a card, a date range, and a recurring amount, they are very likely the same vendor. If they share only a token like "Supply," they probably are not. Machine-readable identifiers are the most reliable evidence when they exist: a tax ID or an exact address outranks a name similarity score.
If you want a starting point rather than a blank column, an Inferred Column can propose a canonical grouping or normalized name for each row. Treat it as a draft. Check it against the spend totals before you rely on it, and expect to correct the long tail by hand. The judgment stays with you because the consequence of a wrong merge (two real suppliers collapsed into one, hiding a relationship you wanted to renegotiate) is worse than the cost of reviewing the candidates.
This is a narrower task than standardizing the formats inside one vendor's invoices, which we cover separately in standardizing vendor invoice data. Here the format is already handled at extraction time; what you are building is the identity layer on top of it.
What This Approach Still Doesn't Do
The steps above produce a clean list. They do not remove the judgment, and it is worth being clear about where the automation stops.
It is not an automatic vendor-master deduplication engine. Extraction gets every record into one sheet with the raw vendor names; it does not decide which names are the same entity. That mapping is yours to define, and the tool's job is to make that mapping fast to build and easy to verify.
Fuzzy matching still fails on unrelated strings such as descriptor codes, so a purely algorithmic pass will leave the AMZN MKTP US type of line unmerged. Whether a parent company and its divisions are one vendor or several is a business decision, not a technical one, and no tool can make it for you. And if a reimbursement carries only the employee's name with no receipt attached, there may be no way to recover the vendor from the record at all. Those cases have to be resolved from the card statement or from the employee, not from the data.
Reconciling the card statement against your books is a different job again, and one where mixing personal and company cards creates its own problems, which we cover in credit card reconciliation. Getting the vendor layer right first makes that reconciliation easier, because you stop re-deciding the same supplier every month.
Deduplicating a Vendor List: FAQ
Can't I just use Excel Power Query fuzzy matching on the vendor list?
You can use it, and it will catch some variants, but it has two blind spots for this specific job. It cannot connect a descriptor code to a brand name, because those are not similar strings. And the threshold that catches suffix variants also merges unrelated names, so you end up reviewing every candidate anyway. Power Query is most useful once the records are already in one sheet and you are tuning the candidate list, not as the first and only step.
Do I need a full vendor master system for a company this size?
For a 100 to 300 person company, a spend-management platform (Ramp, Brex, Expensify, Bill.com) solves the ongoing problem by routing future spend through cards that categorize vendors and flag duplicates at the point of purchase, and tools like Coupa and Zip extend that to larger procurement. None of them rebuild the record of what was already bought on personal cards. That historical list still has to be reconstructed from documents, which is the part this approach handles.
The AI extracted the vendor names, but they are still inconsistent. What now?
That is expected. Extraction reproduces the name as it appears in the source, and the source is inconsistent. The next step is the canonical mapping: sort by spend, group the variants, and assign one canonical name per real supplier. The extracted table with a Vendor (canonical) column is the deliverable, and it is what makes the totals trustworthy.
How do I handle reimbursements that only show the employee's name?
Use the receipt or card statement as the source of truth for the vendor, and join it to the reimbursement by amount and date. If the receipt is missing, the card descriptor is the fallback, which is why capturing the raw vendor string from the statement matters even when it looks like a code. Where both are missing, the expense is not recoverable from the data and has to be chased manually.
How often should this list be rebuilt?
Once you have the canonical mapping, ongoing maintenance is lighter because new purchases from existing suppliers map to names you have already defined. Re-run the extraction and mapping for new spend, and review the unmatched names on a schedule. Quarterly is the cadence AP teams commonly use for vendor master review, and it keeps the list from decaying back into a pile of variations.
The bottleneck in vendor consolidation is the identity layer, and deciding which supplier each record belongs to is what takes the time. Once that layer exists in a spreadsheet, ranking spend and finding consolidation opportunities takes minutes instead of weeks. Once it does not, every report you build on top of the list is only as stable as the names underneath it. For the wider picture of pulling structured data out of expense documents, see the guide to expense report data extraction. If the records span many months, year-end transaction extraction and the credit card reconciliation pipeline show how the same extracted table feeds reconciliation and the three-way matching without an ERP work that follows it.