Customer analytics · LLM · Data integrationCJ AI Center, cross-affiliateWrite-up October 2026 · 6 min read

A common language for customers across businesses

Each business in a group knows its own customers through its own products, categories and IDs. This is how an LLM-built interest taxonomy, an affinity data mart and persona estimation made those customers comparable across businesses — and a small version you can explore.

Built withSQLPythonpandasLLM labellingSciPy testsStreamlitplotlyAG Grid
CommerceKitchen › CookwareLiving › Air careHealth food › SupplementsPets › SuppliesLeisure › Camping Beauty retailSkincare › CreamSkincare › SerumWellness › Inner beautyHair › TreatmentBody › Fragrance StreamingDrama / RomanceNon-fiction / DocuseriesFilm / ThrillerVariety / CookingAnimation / Family Skin healthWellnessDocumentariesDiet & nutritionHome cooking
Three category trees with nothing in common, and one set of interest labels that all of their items can be mapped onto.

One customer, three partial views

Each business in the group holds data on products, customers, content, purchases and service usage — but a single business's data never shows the whole person. A customer who buys barrier-repair skincare in one place may watch wellness documentaries in another and buy probiotics in a third. Those patterns are invisible inside any one company, and the businesses share no product codes or category names that would let you join them.

Design: a shared interest layer

  1. Decide what to connect. Define the direction of integrated analysis and which data — products, customers, content — should be joined, and structure it for dashboard use.
  2. Make metadata comparable. Product and content metadata and unstructured category information were processed and mapped onto a shared set of interests with definitions, using an LLM so that tens of thousands of items could be labelled consistently.
  3. Estimate personas. Each customer's activity rolls up into an interest profile; LLM-based persona estimation turns a cohort's profile into a readable description.
  4. Show it. A dashboard with affiliate-level insights, persona analysis, content feature extraction and product attribute analysis, so marketing and planning teams can pick a cohort in one business and see it in the others.

Live model, computed in your browser on generated customers of three generated businesses. Left: the three category trees share nothing; the labelling step maps each item onto shared interests (the labels shown were produced in advance, standing in for the LLM). Right: a cohort picked in one business, read in the other two, with lift against everyone and a z-test. Switch cohorts or draw new customers. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

Monthly extracts from three affiliates, an LLM pass that labels every item with a shared interest taxonomy, an affinity data mart, and a dashboard for persona and segment analysis.

1 · Extract

Monthly aggregates

Purchases, viewing, subscriptions and demographics from each affiliate, aggregated monthly in SQL.

SQL
2 · Label

One interest taxonomy

Item and content metadata go to an LLM with a 100-interest taxonomy and its definitions; one to five labels per item, anything outside the taxonomy dropped.

LLMtaxonomy
3 · Affinity

Customer × interest mart

Each customer's activity becomes interest shares and within-label percentiles.

pandasparquet
4 · Analyse

Personas and significance

Interest profiles per segment, LLM persona estimation, and tests for what really over-indexes.

SciPyLLM
5 · Dashboard

Self-serve insight

A Streamlit dashboard with login, data grids and charts for marketing and planning teams.

StreamlitplotlyAG Grid
Stack
LayerTechnologyWhat it does here
DataSQL extracts, pandas, parquetMonthly logs from three affiliates joined on customers
LabellingLLM constrained to a 100-interest taxonomy in 11 groupsMakes home-shopping, beauty and streaming items comparable
AffinityInterest share, within-label percentile, min-max scaled scoresCustomer × interest profiles
TestsOne-sample and Welch t-tests, two-proportion z-testsFlags interests, categories and brands that over-index in a segment
DashboardStreamlit, streamlit-authenticator, AG Grid, plotly, word cloudsAffiliate insight, persona, content and product analysis
Core formulation
share(c, l)  = labelled rows of customer c with label l  /  all labelled rows of c
pct(c, l)    = 100 · (N_l − rank_l(c) + 1) / N_l               among customers holding l
score(l | S) = 0.25·MM₁₀(mean score in S) + 0.75·MM₁₀(share of S holding l)
tests        Welch t on scores;   two-proportion z on categories, brands, items
persona(S)   = LLM( top interests of S, demographics of S )
  • Constrained labelling. The model may only answer with labels from the taxonomy, which is what makes tens of thousands of items consistent.
  • Scores the business can read. Mean intensity and penetration are scaled to 0–10 and blended, then filtered by significance.
  • Segments as rules. Segments are stored as conditions and recomputed when a screen opens, so they follow new data.
In production vs in the live model
ComponentIn productionIn the live model above
LabellingLLM over tens of thousands of itemsLabels prepared in advance for 36 generated items
AnalysisDashboard scores with t / z tests per segmentCohort lift with a two-proportion z-test
PersonaLLM-estimated from the profileTemplate sentence from the same profile
DataCustomers of three affiliates4,000 generated customers of three generated businesses

Why labels, not joins

The tempting approach is to map categories to categories. It fails quickly: category trees differ in depth and meaning, and content has no "category" in the retail sense at all. Interests describe the person, not the shelf, which is why the same label can attach to a cream, a humidifier and a documentary. The quality of labelling therefore caps everything after it, and the taxonomy's definitions matter as much as the model.

With labels in place, the analysis moves from product- and content-level descriptive metrics to customer behaviour and persona-based understanding: which interests over-index in a cohort, in which business, and whether the difference is statistically real rather than an artefact of cohort size.

What it changed

Integrateddashboard for cross-affiliate customer insight
PersonaLLM-based estimation from interest profiles
Investmentthe PoC supported a group-level data-platform investment and service launch

The project built an integrated analytics dashboard that enabled cross-affiliate customer insight for the first time, and it made the value case concrete enough to support a group-level data-platform investment and a service launch.

Limitations

About the demo and confidentiality

Businesses, items, categories, labels and customers in the embedded model are generated. No customer, product, schema or business result from the real project appears here.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterIntegration design, data modelling, persona logic, analysis pipeline and visualization structure. Demo re-implemented on synthetic data for this site.