Data Platform Engineering for a Cosmetics Maker

A cosmetics manufacturer sells almost entirely through retailers, so its only view of the end consumer is data scattered across 24 countries. GroupBWT built one platform for it.

unified customer data platform across 24 retail markets

CLIENT STORY

A global decorative-cosmetics manufacturer sells through retail chains in dozens of countries rather than direct to shoppers. Its in-house data team had designed an enterprise warehouse on a three-layer Data Vault 2.0 model and wanted it fed by every signal the business touches — what sells on each retailer’s shelf, how the brands rank against competitors, and how marketing performs across channels. The goal was to turn that data into decisions for its sales and marketing teams worldwide.

Service: Data Engineering
Industry: Beauty & Personal Care
Region: EU

We didn't want to buy a platform and bend our warehouse to fit it. We wanted engineers who would expand it with us, source by source, inside our own stack. — Senior Data Engineer

Our marketing analytics and sales teams are already using the data from these pipelines. Last week there was a request on app and marketplace reviews, and it came straight from the platform — no one had to chase it down. — Team Lead, MarTech

Introduction

The Challenge: Every Retailer a Different File, Every Country a Blind Spot

Selling through retailers means the manufacturer never meets its end customer directly. The only way to see what happens at the shelf is the data each retailer and distributor sends back — and across 24 countries, no two send it the same way. One ships a monthly Excel with separate online and offline tabs; another a weekly CSV; a distributor packs ten countries into one report. On top of that sit syndicated market-share data and half a dozen marketing platforms, several with no API at all.

The client’s data team had the warehouse design and the modeling discipline. What it didn’t have was the engineering capacity to keep up: every new retailer, market, or platform meant another bespoke pipeline, and sources were multiplying faster than the team could integrate them. Until they were connected, each region — Italy, South Africa, Brazil — worked from its own numbers, and the sales team couldn’t get a reliable read on where to extend an assortment.

scattered retailer files blocking cross-country sellout visibility
The Solution

An Embedded Team and a Pipeline Factory Behind One Warehouse

The client didn’t want a platform sold to them — they wanted engineers inside their stack. GroupBWT started with a six-week proof of concept (three retailer pipelines), then grew into a standing data-engineering team working in the client’s own repositories, CI/CD, and review process. What started as three pipelines is now more than 40 in production, pulling from over 24 countries.

Embedded delivery. No code thrown over a wall. The team works inside the client’s own setup — their Git, their deploy pipeline, their task board — and every pipeline passes the client’s review before it ships, exactly the way an in-house engineer’s would.

A sellout pipeline factory. Every retailer’s file arrives in a different shape. The pipeline parses it, maps it onto one common record — country, currency, product identifier, quantity, revenue — and pushes it through the Data Vault’s raw, business, and presentation layers. The real call was building one reusable template up front instead of a script per retailer, which is why a new source takes a day now, not a week.

External-source integration. Shelf data was only the start. The team also folded in syndicated market-share feeds — cleaning up inconsistent competitor brand names so the client could finally rank against rivals — along with paid, organic, and social-listening data from five marketing platforms. A couple of those had no API, so the manual exports got automated instead.

AI-assisted documentation. Documentation usually lags behind delivery. As the pipelines multiplied, the team built a small in-house tool that drafts each one’s docs with an LLM and posts them to the knowledge base. An engineer signs off on every page before it goes live, so the docs stay current instead of rotting.

Tech stack: Python, Polars, PostgreSQL, Azure, Data Vault 2.0, Drone CI, ArgoCD, Grafana, Tableau.

reusable pipeline template normalizing every retailer's file

The unglamorous decision was building a pipeline template before building pipelines. It cost us time up front, but it's why onboarding the twentieth retailer looks nothing like the first — the format changes, the model doesn't, and a new source is a day's work instead of a small project.

Alex Yudin
Alex Yudin
Head of Data Engineering, GroupBWT
The Results

24+ Countries of Retail Data on One Platform, a New Source in a Day

  • By normalizing every retailer’s file to one model in the warehouse, GroupBWT gave the client a single view of shelf sales across 24+ countries instead of 24 incompatible reports.
  • Onboarding a new retailer or market now takes about a development day, so the warehouse keeps up even as sources pile on.
  • With syndicated market-share data and five marketing platforms in one model, retail, competitive, and marketing signals finally sit side by side.
  • The sales and marketing teams now pull assortment and review signals straight from the platform, turning scattered country data into decisions about where to extend a range.
  • Working inside the client’s own repositories and deployment flow let GroupBWT scale from a six-week pilot to a standing team without the client adopting a new tool.
24+
Countries Unified
~1 day
To Onboard a New Source
40+
Pipelines in Production
global retail platform driving assortment and marketing decisions

Selling through retailers but blind to the shelf?

If your business runs through retail partners and your customer data is scattered across countries and file formats, we map it, model it, and build the pipelines — starting with a proof of concept on your real data.

Contact Us