Retail Catalog Matching at Scale: AI Product Matching Without a Shared Barcode
This case explains how GroupBWT designs AI product matching, competitor catalog comparison, and PDP monitoring engagements — the architecture we deploy for catalog matching projects, with components already running in production for retail and CPG accounts.
CLIENT STORY
A catalog operations team at a top-10 US mass retailer came to us looking for a way out of the back-room loop — a small contractor crew lifting each item off the shelf, measuring it against the website, putting it back. The online-only assortment, items that never sit on a shelf, was unreachable by definition.
The harder problem surfaced halfway through the call. The same household product — a laundry detergent, a roll of toilet paper, a hand lotion — sells at three retailers under three different UPCs, three titles, three attribute schemas. Their own-brand assortment, a strategic growth lever, had no equivalent anywhere to anchor against. Manual reconciliation could neither cover the digital shelf nor solve the matching question.
That conversation crystallized two design choices we now bring to catalog-matching engagement: split exact-match and like-item from day one, and treat own-brand SKUs as a distinct retrieval problem with their own metrics — not a leftover bucket.
| Service: | Web Scraping + Data Engineering + Data Science |
|---|---|
| Industry: | Retail / CPG |
| Region: | North America |
Read summarized version with
A back-room crew lifts each item, checks size and weight against the website, then puts it back. Online-only SKUs are unreachable. The harder problem is cross-retailer matching — the same item carries a different UPC at every retailer, and own-brand SKUs have no equivalent to anchor against. — Director, US retailer
Item data lands from vendors, syndicators, and internal aggregators across millions of SKUs, and category requirements diverge — apparel needs colorway accuracy, electronics needs the right product video, grocery needs nutritional fields. Item-by-item review does not scale; we need a way to surface gaps in mass. — Head of Data, US retailer
Catalog Quality Without a Shared Identifier
Off-the-shelf matching APIs return one similarity score per pair with no decision context — unsafe to wire directly into a PIM workflow. The brief from a catalog operations team usually arrives in three parts:
- Cover the entire assortment, including online-only SKUs, on a monthly cycle.
- Match exact equivalents and like-items separately — a category lead must never see the two confused in one screen.
- Give analysts a queue with red flags and full attribute context, not a black-box probability.
A Three-Layer Pipeline for AI Product Matching and PDP Monitoring
Layer 1 — Cross-Retailer PDP Extraction. Targeted crawls of Amazon, Walmart, and peer US retailers extract full PDPs — attributes, images, video, review signals — normalized to one attribute schema.
Layer 2 — Two-Mode Matcher. Exact match returns the same SKU at a competitor; like-item returns the closest substitute by purchase-driving attributes, with a confidence score and a short reason. The two modes are stored and shown separately. Matching models are tuned per category, with confidence thresholds adjusted to category-specific error tolerance. Low-confidence pairs are routed into an analyst review queue, and analyst decisions feed the next retraining cycle. We recommend reporting accuracy on held-out data per category, not one headline number.
The matcher is the visible part, but what decides success is upstream — clean PDP extraction, schema normalization, and a category-level evaluation set the team trusts. We split exact match and like-item on day one, because conflating them costs more to unwind later than to design correctly now.
Analyst Comparison Dashboard
Layer 3 — Analyst Comparison Dashboard. One screen per product: PIM values left, matched competitor values right, gaps flagged, confidence visible. The analyst confirms or dismisses each gap in one click. Output: a daily “red list” of catalog gaps against an agreed SLA.
Reference stack: Scrapy with hybrid proxy rotation; Postgres + S3; sentence-embedding models fine-tuned where category data volume justifies it; analyst feedback loop; Metabase; Kubernetes.
What This Pipeline Delivers, and Why
Three decisions separate this from a generic data-science engagement.
Retail-specific data engineering, not just modeling.
The matcher works only because the upstream extractor produces a clean, schema-aligned dataset. Crawl, normalization, matching, and dashboarding stay in one team — no integration gaps between vendors.
Two modes by design, not by patch.
Exact match and like-item are split from day one. A category review never accidentally compares an own-brand item against an unrelated competitor SKU.
Analyst-first dashboard, model-second.
The screen the catalog team works in is the deliverable. Teams adopt it because it makes an existing job faster — not because they trust an opaque score.
We report four metrics at pilot close and monthly thereafter:
- Exact-match F1 on a held-out category sample
- Like-item top-3 hit rate per own-brand segment
- Analyst queue clearance time against the agreed SLA
- Cycle time from competitor PDP change to PIM flag
- Benchmarks from production engagements are shared under NDA at scoping.
Looking to Automate Retail Catalog Matching Across Amazon, Walmart, and Peer Retailers?
GroupBWT designs AI product matching, competitor catalog comparison, and PDP monitoring engagements for retailers and CPG brands without shared UPCs, at scope from a few hundred SKUs to multi-million-SKU catalogs.
You have an idea?
We handle all the rest.
How can we help you?