A common language for customers across businesses
Each business in a group knows its own customers through its own products, categories and IDs. This is how an LLM-built interest taxonomy, an affinity data mart and persona estimation made those customers comparable across businesses — and a small version you can explore.
One customer, three partial views
Each business in the group holds data on products, customers, content, purchases and service usage — but a single business's data never shows the whole person. A customer who buys barrier-repair skincare in one place may watch wellness documentaries in another and buy probiotics in a third. Those patterns are invisible inside any one company, and the businesses share no product codes or category names that would let you join them.
Design: a shared interest layer
- Decide what to connect. Define the direction of integrated analysis and which data — products, customers, content — should be joined, and structure it for dashboard use.
- Make metadata comparable. Product and content metadata and unstructured category information were processed and mapped onto a shared set of interests with definitions, using an LLM so that tens of thousands of items could be labelled consistently.
- Estimate personas. Each customer's activity rolls up into an interest profile; LLM-based persona estimation turns a cohort's profile into a readable description.
- Show it. A dashboard with affiliate-level insights, persona analysis, content feature extraction and product attribute analysis, so marketing and planning teams can pick a cohort in one business and see it in the others.
Live model, computed in your browser on generated customers of three generated businesses. Left: the three category trees share nothing; the labelling step maps each item onto shared interests (the labels shown were produced in advance, standing in for the LLM). Right: a cohort picked in one business, read in the other two, with lift against everyone and a z-test. Switch cohorts or draw new customers. Open the live model on its own page ↗
Architecture, stack and core formulation
Monthly extracts from three affiliates, an LLM pass that labels every item with a shared interest taxonomy, an affinity data mart, and a dashboard for persona and segment analysis.
Monthly aggregates
Purchases, viewing, subscriptions and demographics from each affiliate, aggregated monthly in SQL.
One interest taxonomy
Item and content metadata go to an LLM with a 100-interest taxonomy and its definitions; one to five labels per item, anything outside the taxonomy dropped.
Customer × interest mart
Each customer's activity becomes interest shares and within-label percentiles.
Personas and significance
Interest profiles per segment, LLM persona estimation, and tests for what really over-indexes.
Self-serve insight
A Streamlit dashboard with login, data grids and charts for marketing and planning teams.
| Layer | Technology | What it does here |
|---|---|---|
| Data | SQL extracts, pandas, parquet | Monthly logs from three affiliates joined on customers |
| Labelling | LLM constrained to a 100-interest taxonomy in 11 groups | Makes home-shopping, beauty and streaming items comparable |
| Affinity | Interest share, within-label percentile, min-max scaled scores | Customer × interest profiles |
| Tests | One-sample and Welch t-tests, two-proportion z-tests | Flags interests, categories and brands that over-index in a segment |
| Dashboard | Streamlit, streamlit-authenticator, AG Grid, plotly, word clouds | Affiliate insight, persona, content and product analysis |
share(c, l) = labelled rows of customer c with label l / all labelled rows of c pct(c, l) = 100 · (N_l − rank_l(c) + 1) / N_l among customers holding l score(l | S) = 0.25·MM₁₀(mean score in S) + 0.75·MM₁₀(share of S holding l) tests Welch t on scores; two-proportion z on categories, brands, items persona(S) = LLM( top interests of S, demographics of S )
- Constrained labelling. The model may only answer with labels from the taxonomy, which is what makes tens of thousands of items consistent.
- Scores the business can read. Mean intensity and penetration are scaled to 0–10 and blended, then filtered by significance.
- Segments as rules. Segments are stored as conditions and recomputed when a screen opens, so they follow new data.
| Component | In production | In the live model above |
|---|---|---|
| Labelling | LLM over tens of thousands of items | Labels prepared in advance for 36 generated items |
| Analysis | Dashboard scores with t / z tests per segment | Cohort lift with a two-proportion z-test |
| Persona | LLM-estimated from the profile | Template sentence from the same profile |
| Data | Customers of three affiliates | 4,000 generated customers of three generated businesses |
Why labels, not joins
The tempting approach is to map categories to categories. It fails quickly: category trees differ in depth and meaning, and content has no "category" in the retail sense at all. Interests describe the person, not the shelf, which is why the same label can attach to a cream, a humidifier and a documentary. The quality of labelling therefore caps everything after it, and the taxonomy's definitions matter as much as the model.
With labels in place, the analysis moves from product- and content-level descriptive metrics to customer behaviour and persona-based understanding: which interests over-index in a cohort, in which business, and whether the difference is statistically real rather than an artefact of cohort size.
What it changed
The project built an integrated analytics dashboard that enabled cross-affiliate customer insight for the first time, and it made the value case concrete enough to support a group-level data-platform investment and a service launch.
Limitations
- Insight is only as good as the labelling; ambiguous items and thin metadata produce noisy interest profiles.
- Cross-business comparison needs customers who appear in more than one business, and that overlap is uneven.
- In the live model the item labels are prepared in advance and the persona text is templated; production used an LLM for both. The numbers come from a generator that plants interest structure, so they say nothing about real customers.
About the demo and confidentiality
Businesses, items, categories, labels and customers in the embedded model are generated. No customer, product, schema or business result from the real project appears here.