CLIENT STORY
A UK fashion-tech startup is building two products on one data foundation: a consumer price-comparison platform across resale marketplaces, and a B2B tool that turns a product photo into material, pricing history, and source matches. Both depend on the same fifty-million-record catalog. One vendor going down meant both products stalled in the same week.
| Service: | Web Scraping |
|---|---|
| Industry: | E-Commerce |
| Year: | 2024–2026 |
| Location: | EU (UK) |
Read summarized version with
We needed someone to manage scraping for us — set expectations, hit the schedule, keep data flowing. GroupBWT delivered, week after week, for nearly two years. - Data Insights Lead, UK Fashion-Tech Startup
We came in with a single API and a worried CTO; we now have six live scrapers on a schedule we trust. - Co-Founder, UK Fashion-Tech Startup
The Challenge: One Vendor, Fifty Million Records, And No Plan B
Fifty million product records sat in the warehouse — every one from a single third-party API. The CTO tried to build the first scraper himself; it held up for an afternoon before the marketplace updated its anti-bot stack. Scraping wasn’t core work, and investors wanted a diversification plan before the funding review.
Every alternative source sat behind PerimeterX, CloudFlare, or per-query API caps. TheRealReal stopped at 17 pages of pagination. Vestiaire Collective’s API capped queries at 10,000 results. eBay’s bidding metadata defied any clean schema. Residential proxy traffic showed up on a finance line the founders had to defend.
Buying more vendor APIs would not close the gap: feeds skip long-tail SKUs, hold no history, and reprice without notice. If the single source ever shut down, both product lines stalled in the same week.
Six-Marketplace Scraping Factory With Hybrid Proxy Strategy
The platform went from one vendor API to six independent scrapers — brought online one at a time across the first year, each with site-specific anti-bot tuning on the same operational pattern.
Pilot on TheRealReal: PerimeterX as the Hardest Test
The client started with one scoped scraper at a fixed price — no infrastructure to stand up, no long contract — so TheRealReal had to prove the model before the rest were greenlit. Anti-bot controls cleared ~390K products in session one; retuned pacing cut the run from seven days to two — a template ported to every later source.
Hybrid Proxy Routing: Cheap Pool by Default, Premium Calls When Needed
Catalog runs on commodity residential pools; ~1 in 400 requests routes through a paid Unlocker — the minimum premium share that keeps the session valid. Images moved to flat-fee Storm Proxies at $147/month. The same pattern runs across all six sources.
On scraping, the win isn't beating anti-bot once. It's finding the cheapest traffic mix that keeps the session valid month after month.
Vinted at 40M Records: Why the Queue Broke and How We Rebuilt It
Vinted’s first run broke the queue — 300+ duplicates entering every second, faster than the database could handle. A pre-queue deduplication layer screened 40M product IDs before they reached the database — and the problem disappeared.
Operations Layer: Kubernetes, Dashboards, and Maintenance the Client Doesn’t Touch
Images sit in cold-tier S3 storage — archive pricing, millisecond reads when the B2B tool needs them. Metabase tracks every session on a dashboard the client can see. Maintenance is GroupBWT’s: when StockX or Vinted update their site structure, the client writes zero code and books zero hours on anti-bot reactivation.
Tech stack: Scrapy on Kubernetes, RabbitMQ + Redis, MySQL + AWS S3, Metabase; Oxylabs (catalog residential pool), Storm Proxies (flat-fee image traffic), BrightData Unlocker (session refresh).
Six-Marketplace Data Pipeline
- The session time dropped from seven days to two after pacing and proxy retuning, freeing weekly capacity for new sources.
- ~5.5M-product catalog ingested every month, despite a 10,000-result API cap.
- Monthly maintenance keeps every scraper alive without client engineering — when sources update their structure, GroupBWT restores in-window.
- Cold-tier image storage keeps lookups fast for the B2B tool while pulling the storage line down.
- Twenty consecutive months of cleared invoices on a per-scraper pricing model — predictable cost on both sides.
Single Data Vendor, Single Point of Failure?
If your platform runs on one API, GroupBWT scopes a fixed-price pilot scraper — tested against anti-bot protection, live within weeks — before the next funding or pricing review.
You have an idea?
We handle all the rest.
How can we help you?