Who are the true fans? An engagement score and rank
The business wanted to find its true fans and make more of them. Everyone agreed such customers existed; no single number — usage, subscriptions, recent activity — identified them. This is how an intuition became an operational score: define engagement by what it should predict, measure each service's recency, frequency and intensity, cut those into levels, and let the data say which services make a fan.
An intuition that needed a definition
The business wanted to define true fans or highly engaged customers in a data-driven way, so that it could find them, count them and grow their number. Single indicators — service usage, subscription count, recent activity — were insufficient to capture engagement and long-term customer value: the heaviest data users were not the most loyal, and the longest-tenured were not always the most active. Before any model, the concept needed a definition the business would accept.
What was built
- Define engagement by its consequences. Engagement was defined in relation to ARPU and churn rate: an engaged customer pays more and leaves less. That turns a feeling into a target a model can be fitted to.
- Describe usage with RFM. For each service, recency, frequency and monetary-style intensity variables were created, and the relationship between usage behaviour and customer value was analysed.
- Cut usage into levels. Raw usage is skewed and non-linear. Optimal binning discretized each variable into levels at the cut points that best separate the target; a genetic algorithm searched the cut points.
- Weight the services. Least-squares regression on the levelled variables derived the weight of each service and each facet, and from them a score and a rank for every customer.
Live model, computed in your browser on generated customers of six services. Left: customers ranked by raw usage count, and how churn and revenue fall across the deciles of that ranking on holdout customers. Right: one variable's levels as found by the genetic algorithm, the service weights, and the same deciles for the engagement score. Switch to equal-count bins to see what the optimization adds. Open the live model on its own page ↗
Architecture, stack and core formulation
Engagement defined against ARPU and churn, measured through recency, frequency and intensity of each service, levelled by optimal binning and weighted by regression into a score and a rank.
Define the fan
Engagement defined in relation to ARPU and churn: an engaged customer pays more and leaves less.
Usage per service
Recency, frequency and intensity variables for each service.
Levels, not raw counts
Optimal binning with a genetic algorithm finds cut points that best separate the target.
Regression
Least-squares (and PLS) regression on the levelled variables gives service weights and a customer score.
Score and rank
A score and rank for every customer, used for campaigns and service priorities.
| Layer | Technology | What it does here |
|---|---|---|
| Target | Combination of ARPU and churn | Makes 'true fan' measurable |
| Features | RFM variables per service | Recency, frequency, intensity of usage |
| Discretization | Optimal binning; genetic-algorithm search over cut points | Non-linear, thresholded usage effects |
| Weights | Least-squares / partial least squares regression | Service and facet weights, a score per customer |
| Related | Rank aggregation (Borda, hierarchical) for LG Hausys targeting | Combining rankings from sources without a common scale |
target_i = z(ARPU_i) − λ·churn_i (live model's choice of λ)
cut points c* = argmax_c η²( target | bin(x; c) ) searched by a genetic algorithm
η² = between-level variance / total variance
score_i = Σ_(service s) Σ_(facet f ∈ {R, F, M}) β_sf · v_sf( level_sf,i )
v = target mean of the level; β by least squares (PLS for collinear facets)
Borda score(i) = Σ_(source k) ( N − rank_k(i) ) rank aggregation- The target does the defining. Agreeing that engagement means higher revenue and lower churn made every later choice testable.
- Levels are explainable. Each level has a plain meaning and a contribution the business can read; the weights say which services are worth growing.
- GA when the threshold matters. A genetic search finds the threshold the data knows and the analyst does not — at small cost where effects are smooth.
| Component | In production | In the live model above |
|---|---|---|
| Binning | Optimal binning with a genetic algorithm | The same: GA over quantile candidates, or equal-count bins |
| Weights | Least squares / PLS | Ridge least squares on level means |
| Data | Real usage, ARPU and churn | 5,000 generated customers of six services |
Design notes
The target does the defining
Most of the project's value came before the modelling: agreeing that engagement means higher revenue and lower churn, in that combination, made every later choice testable. A score that merely counted activity would have rewarded the service everyone uses daily and said nothing about loyalty.
Levels make the score explainable
Binning costs a little precision and buys a great deal: each level has a plain meaning ("watches IPTV at least twice a week"), each level's contribution is a number the business can read, and the weights say which services are worth growing. The genetic algorithm matters where a variable has a threshold the data knows and the analyst does not.
A cousin: combining ranks instead of scores
A related target-marketing project for LG Hausys — interest indices for marriage, moving and interior remodelling, built from web and app activity and purchase records — faced a different version of the same question: several signal sources each gave a ranking, none agreed, and no common scale existed. There, rank aggregation methods such as Borda counts combined the rankings directly, which is the right tool when sources cannot be put on one scale and the wrong one when, as with engagement, a target exists to calibrate against.
Outcome
The project converted an intuitive concept of fans into an operational Engagement Score and Rank. It supported customer segmentation, service-usage campaigns and retention strategies, and the prioritisation of services that contribute to customer value.
Limitations
- The score inherits its definition: a different weighting of revenue against churn produces a different ranking, and that choice belongs to the business.
- Levels fitted on one period drift as services and plans change; cut points and weights need periodic refitting.
- In the live model the engagement structure is planted and smooth, which flatters both the usage count and the score; the gain from optimized binning there is modest.
About the demo and confidentiality
Customers, services, usage, revenue and churn in the embedded model are generated. No usage, billing or churn data, service list or weight from the real project appears here.