Customer analytics · Scoring · OptimizationLG Uplus, services and marketingWrite-up October 2026 · 6 min read

Who are the true fans? An engagement score and rank

The business wanted to find its true fans and make more of them. Everyone agreed such customers existed; no single number — usage, subscriptions, recent activity — identified them. This is how an intuition became an operational score: define engagement by what it should predict, measure each service's recency, frequency and intensity, cut those into levels, and let the data say which services make a fan.

Built withPythonRFMOptimal binningGenetic algorithmLeast squares · PLSRank aggregation
Mobile dataIPTVStreaming appMusicPayments appMembership store × 0.06× 0.31× 0.14× 0.08× 0.27× 0.14 ENGAGEMENT82 / 100 orange ticks = level cut points found by optimization · weights from least squares against revenue and churn
Each service's usage is cut into levels at data-chosen points; each level carries what it predicts; weights per service combine them into one score. Illustrative values.

An intuition that needed a definition

The business wanted to define true fans or highly engaged customers in a data-driven way, so that it could find them, count them and grow their number. Single indicators — service usage, subscription count, recent activity — were insufficient to capture engagement and long-term customer value: the heaviest data users were not the most loyal, and the longest-tenured were not always the most active. Before any model, the concept needed a definition the business would accept.

What was built

  1. Define engagement by its consequences. Engagement was defined in relation to ARPU and churn rate: an engaged customer pays more and leaves less. That turns a feeling into a target a model can be fitted to.
  2. Describe usage with RFM. For each service, recency, frequency and monetary-style intensity variables were created, and the relationship between usage behaviour and customer value was analysed.
  3. Cut usage into levels. Raw usage is skewed and non-linear. Optimal binning discretized each variable into levels at the cut points that best separate the target; a genetic algorithm searched the cut points.
  4. Weight the services. Least-squares regression on the levelled variables derived the weight of each service and each facet, and from them a score and a rank for every customer.

Live model, computed in your browser on generated customers of six services. Left: customers ranked by raw usage count, and how churn and revenue fall across the deciles of that ranking on holdout customers. Right: one variable's levels as found by the genetic algorithm, the service weights, and the same deciles for the engagement score. Switch to equal-count bins to see what the optimization adds. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

Engagement defined against ARPU and churn, measured through recency, frequency and intensity of each service, levelled by optimal binning and weighted by regression into a score and a rank.

1 · Target

Define the fan

Engagement defined in relation to ARPU and churn: an engaged customer pays more and leaves less.

target design
2 · RFM

Usage per service

Recency, frequency and intensity variables for each service.

RFM
3 · Binning

Levels, not raw counts

Optimal binning with a genetic algorithm finds cut points that best separate the target.

optimal binningGA
4 · Weights

Regression

Least-squares (and PLS) regression on the levelled variables gives service weights and a customer score.

least squaresPLS
5 · Rank

Score and rank

A score and rank for every customer, used for campaigns and service priorities.

ranking
Stack
LayerTechnologyWhat it does here
TargetCombination of ARPU and churnMakes 'true fan' measurable
FeaturesRFM variables per serviceRecency, frequency, intensity of usage
DiscretizationOptimal binning; genetic-algorithm search over cut pointsNon-linear, thresholded usage effects
WeightsLeast-squares / partial least squares regressionService and facet weights, a score per customer
RelatedRank aggregation (Borda, hierarchical) for LG Hausys targetingCombining rankings from sources without a common scale
Core formulation
target_i     = z(ARPU_i) − λ·churn_i                         (live model's choice of λ)

cut points   c* = argmax_c  η²( target | bin(x; c) )         searched by a genetic algorithm
             η² = between-level variance / total variance

score_i      = Σ_(service s) Σ_(facet f ∈ {R, F, M})  β_sf · v_sf( level_sf,i )
             v = target mean of the level;   β by least squares (PLS for collinear facets)

Borda        score(i) = Σ_(source k) ( N − rank_k(i) )         rank aggregation
  • The target does the defining. Agreeing that engagement means higher revenue and lower churn made every later choice testable.
  • Levels are explainable. Each level has a plain meaning and a contribution the business can read; the weights say which services are worth growing.
  • GA when the threshold matters. A genetic search finds the threshold the data knows and the analyst does not — at small cost where effects are smooth.
In production vs in the live model
ComponentIn productionIn the live model above
BinningOptimal binning with a genetic algorithmThe same: GA over quantile candidates, or equal-count bins
WeightsLeast squares / PLSRidge least squares on level means
DataReal usage, ARPU and churn5,000 generated customers of six services

Design notes

The target does the defining

Most of the project's value came before the modelling: agreeing that engagement means higher revenue and lower churn, in that combination, made every later choice testable. A score that merely counted activity would have rewarded the service everyone uses daily and said nothing about loyalty.

Levels make the score explainable

Binning costs a little precision and buys a great deal: each level has a plain meaning ("watches IPTV at least twice a week"), each level's contribution is a number the business can read, and the weights say which services are worth growing. The genetic algorithm matters where a variable has a threshold the data knows and the analyst does not.

A cousin: combining ranks instead of scores

A related target-marketing project for LG Hausys — interest indices for marriage, moving and interior remodelling, built from web and app activity and purchase records — faced a different version of the same question: several signal sources each gave a ranking, none agreed, and no common scale existed. There, rank aggregation methods such as Borda counts combined the rankings directly, which is the right tool when sources cannot be put on one scale and the wrong one when, as with engagement, a target exists to calibrate against.

Outcome

Score & rankfor every customer, from an idea that had no definition
Service weightsshowing which usage contributes to value
Campaignsfor usage conversion, usage growth and retention, prioritised by the score

The project converted an intuitive concept of fans into an operational Engagement Score and Rank. It supported customer segmentation, service-usage campaigns and retention strategies, and the prioritisation of services that contribute to customer value.

Limitations

About the demo and confidentiality

Customers, services, usage, revenue and churn in the embedded model are generated. No usage, billing or churn data, service list or weight from the real project appears here.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Target definition, RFM design, binning and weighting, score construction. Demo re-implemented on generated data for this site.