Ten thousand rows, none of them usable as they are
A new industrial supplier sends their catalogue: a spreadsheet exported from their ERP, with columns named in their own shorthand, dimensions in millimetres for some lines and inches for others, pack sizes buried in the description, and no images for half the range. Another supplier sends a PDF price list. A third sends a feed that repeats every variant as its own product.
Your catalogue team spends days per supplier reshaping data before a single product can go live, and buyers still complain that listings are missing the specifications they need to order.
Why supplier data arrives in such poor shape
Suppliers' product data was built for their own purposes: their ERP, their price lists, their reps. It is not structured the way your marketplace categories and filters need. Trade buyers search and filter on specific attributes, such as material, size, thread, voltage, pack quantity or standard, and those are often missing or hidden in free text.
Every supplier is different, so a one-off cleanup does not help with the next one, and suppliers update their data regularly, bringing the same problems back.
| Data problem | Effect on buyers |
|---|---|
| Missing key attributes | Products do not appear in filtered searches |
| Mixed units | Wrong comparisons, wrong orders |
| Pack size in description only | Price per unit misunderstood |
| Variants as separate products | Cluttered search results |
| Poor or no images | Buyers do not trust the listing |
What poor catalogue data costs
Products buyers cannot find are products they cannot buy. Products with unclear pack sizes or units lead to wrong orders, returns and disputes between buyer and supplier, which your team then mediates. Slow catalogue onboarding delays new suppliers going live. And buyers who find listings unreliable go back to phoning suppliers directly.
Search and advertising suffer too. Listings without structured attributes rank poorly in your own search and give little to work with in product feeds or category pages, so good suppliers look worse than they are.
The catalogue intake we build
- Column mapping: each supplier's file layout is mapped to your product schema once, with AI suggesting mappings for new files and a person confirming; the mapping is reused for every update from that supplier.
- Attribute extraction: attributes hidden in descriptions, such as size, material and pack quantity, are extracted into structured fields, with confidence scores.
- Standardisation: units are converted to your standard, values are normalised, such as consistent material names, and variants are grouped under parent products.
- Gap filling: where suppliers provide datasheets or spec sheets, missing attributes are read from those documents, marked as extracted so they can be checked.
- Completeness scoring: every listing gets a score against the attributes your category requires, and listings below your threshold are held or shown with lower prominence, as you decide.
- Supplier feedback: suppliers receive a report of specific fixes, such as 40 products missing voltage, with a template to return, rather than a general request to improve their data.
What catalogue work looks like afterwards
New supplier catalogues go through the intake, and the catalogue team reviews the uncertain extractions and mappings rather than reshaping every row. Updates from existing suppliers are processed with the saved mapping. Buyers see consistent attributes and filters that work across suppliers, and your team can see which suppliers' data is holding back their sales.
Is your catalogue data holding you back?
- Every new supplier's data is cleaned by hand.
- Filters return incomplete results because attributes are missing.
- Units and pack sizes are inconsistent between suppliers.
- Buyers order the wrong quantity because pack sizes are unclear.
- Supplier updates reintroduce problems you already fixed.