Causal inference · Time series · Customer retentionLG Uplus, mobile businessWrite-up October 2026 · 7 min read

Finding what leads churn, not what moves with it

A churn score says which customers might leave. It cannot say why the whole company's churn rose this quarter, or what to do about it next quarter. This is the project that structured about four hundred market, competitor, service and customer-experience variables and used causal discovery and inference to separate the indicators that lead churn from those that merely move with it.

Built withPythonpandasPCMCI · tigramitePC algorithmBayesian networksMMHC · Tabu searchPSMIPTWG-formula
t − 3 wkt − 2 wkt − 1 wkthis week Churn rateCompetitor promotionsNetwork complaintsHandset launchesApp store ratingNumber-porting volume orange: leads churn · grey: follows it
A lagged causal graph: competitor promotions lead churn by a week, network complaints by two, handset launches by three — while app ratings and porting volume follow it.

From customer scores to company-level drivers

Churn management is a core business problem for telecommunications and subscription companies, and customer-level churn scores are its standard tool: they identify who is likely to leave and feed retention campaigns. What they do not explain is how market conditions, competitor actions, service changes and customer-experience indicators move company-wide churn — the number that strategy is set against. Business teams had hunches about the drivers, and correlations to back every hunch; what they lacked was a way to tell leading factors from coincident ones.

What was built

  1. Structure the candidates. About 400 potential churn-related variables were collected and structured: external market data, competitor indicators, internal business metrics and customer-experience data, aligned as time series.
  2. Discover the structure. Because the question was about leading factors rather than correlations, time-series causal discovery methods were evaluated — the PC algorithm, PCMCI, Bayesian-network search and Tabu search — to find which variables influence churn, at which lag.
  3. Estimate the effects. With candidate causes identified, propensity score matching, inverse probability of treatment weighting and the G-formula were used to estimate how much each factor moves churn, giving business teams interpretable priorities rather than a ranked correlation table.

Live model, computed in your browser on generated weekly indicators with a planted lagged structure, autocorrelation, a shared drift and a feedback loop in which marketing spend follows churn. Left: the indicators ranked by correlation with churn, tagged afterwards with what the causal analysis found. Right: the lagged graph discovered by PCMCI with partial-correlation tests, and each lead's effect on churn against the planted truth. Try two, three or five years of history. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

About four hundred weekly indicators, several causal-discovery algorithms compared on them, and effect estimation on the drivers they agree on.

1 · Candidates

~400 weekly series

External market data, competitor indicators, internal business metrics and customer-experience data aligned to one weekly grain.

Pythonpandas
2 · Discovery

Lagged causal graphs

PCMCI for time-lagged structure, compared with the PC algorithm and score-based Bayesian-network search (MMHC, Tabu).

tigramitePCMMHC
3 · Effects

How much each driver moves churn

Propensity score matching, inverse probability of treatment weighting and the G-formula on the identified drivers.

PSMIPTWG-formula
4 · Priorities

Lever, lag, size

Drivers ranked by effect and lead time for policy and marketing teams.

reporting
Stack
LayerTechnologyWhat it does here
DataPython, pandasWeekly panel of market, competitor, internal and experience indicators
Time-series discoveryPCMCI (tigramite) with conditional-independence testsLagged parents of churn, free of autocorrelation and common drivers
Structure searchPC algorithm; Bayesian networks by Max-Min Hill-Climbing and Tabu searchCross-checks of the discovered structure
Effect estimationPSM, IPTW, G-formula (outcome modelling)Interpretable effect sizes for each driver
Research toolWeb MVP for time-series causal discovery (Cloud Run)Upload data, choose lags and test strength, read the graph and effects
Core formulation
PCMCI     X^i_(t−τ) → X^j_t   kept iff   X^i_(t−τ) ⊥̸ X^j_t  |  P(X^j_t) \ {X^i_(t−τ)} ∪ P_τ(X^i_(t−τ))
          τ = 1 … τ_max;   CI test = partial correlation in the live model

G-formula E[Y^a] = Σ_l E[Y | A = a, L = l] · P(L = l)
IPTW      ATE = E[ A·Y / e(L) ] − E[ (1 − A)·Y / (1 − e(L)) ],     e(L) = P(A = 1 | L)
PSM       match treated and control units on e(L); compare outcomes
  • Two steps against two traps. The PC step removes most spurious lags; the momentary conditional-independence test conditions on both variables' parents, removing autocorrelation and common-driver links.
  • Algorithms compared, not trusted alone. Constraint-based and score-based searches were run side by side; drivers they agree on carry more weight.
  • From graph to action. Effects were estimated with matching, weighting and outcome models so business teams got a size per lever, not just an arrow.
In production vs in the live model
ComponentIn productionIn the live model above
Variables~400 weekly indicators9 generated indicators with a planted structure
DiscoveryPCMCI, PC, MMHC / Tabu Bayesian networksPCMCI with partial-correlation tests, written in JavaScript
EffectsPSM, IPTW, G-formulaRegression on the discovered parents (linear outcome model)

Design notes

Time changes the question

Cross-sectional causal discovery asks which variables are related given the others. With time series the question gains a direction and a delay: a cause must precede its effect, and it may act through several weeks. Methods built for this — PCMCI in particular — test each lagged candidate while conditioning on the parents of both variables, which strips out the two classic false positives of correlation analysis: a variable that is autocorrelated with a real cause, and a variable that is a consequence of churn itself.

Feedback is the normal case

Retention spend goes up when churn goes up; call-centre waits lengthen when complaints rise; number-porting volume reflects churn that has already happened. In a correlation ranking these look like drivers. Treated as lagged, conditioned relationships they sort themselves out, and the spend that looked like it raised churn turns out to lower it once its timing is accounted for.

From a graph to a decision

A discovered graph is an input, not an answer. The effect sizes — estimated with matching, weighting and outcome models — are what let a business team compare a week of competitor promotion against a point of network complaints and decide where the next retention budget goes.

Outcome

≈400variables structured as candidate drivers
Leading factorswith lags and effect sizes, in place of a correlation table
Churn-rate gapagainst the leading competitor narrowed after the resulting actions

The analysis supported policy and marketing strategies for proactive churn management and contributed to reducing the churn-rate gap against the leading competitor. More durably, it expanded churn management from customer-level prediction into company-level leading-factor management.

Limitations

About the demo and confidentiality

The indicators, their names, the planted relationships and all numbers in the embedded model are invented. No market, competitor, operational or customer data from any operator appears here, and the real variable set is not described.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Variable structuring, method evaluation, causal discovery and effect estimation. Demo re-implemented on generated data for this site.