Experimentation · Statistics · Product analyticsLG Uplus, IPTV and mobile TVWrite-up October 2026 · 6 min read

An experimentation platform for IPTV and mobile TV

New features, screens, recommendation logic and marketing messages were being judged by comparing the weeks after a launch with the weeks before — a method that cannot tell a good change from a good week. This is the platform designed to replace it: a repeatable way to plan, randomize, test and read experiments, built for product teams rather than statisticians.

Built withPythonPower analysisRandomizationTwo-proportion testsBayesian A/BThompson sampling
LAUNCH, THEN COMPARE promo week reports +0.6 pp, "significant" — true effect 0 RANDOMIZE, PLAN, DECIDE planned sample estimates +0.0 pp, CI −0.3 to +0.3 — do not ship
The same month, read two ways: a launch confounded by a promotion week, and a randomized test in which both arms live through the same week and the interval narrows onto the truth. Illustrative numbers.

Good change, or good week?

IPTV and mobile TV services needed a repeatable experimentation platform to evaluate new features, screens, recommendation logic and marketing messages. Without standardized experimental design, results could be biased or confused with natural fluctuation: a new home screen launched in the same fortnight as a content promotion gets credit for the promotion, and a change launched into a quiet week gets blamed for the quiet. Teams were making real decisions on comparisons that could not, in principle, answer the question.

What was built

The platform covers the whole life of an experiment, in the order a product team actually meets it:

Live simulation in your browser. One month of sessions with a weekend lift, drift and a content promotion the analyst does not know about. Left: the new screen launched to everyone on day 15 and judged before versus after. Right: the platform — planned sample, daily randomization, the cumulative lift with its 95% band, a decision at the planned point, the Bayesian probability that B wins, and what Thompson sampling would have done with the same traffic. Set B's true effect and draw a new month. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

The statistical core of an experimentation platform for IPTV and mobile TV: design, assignment, analysis and reporting, with Bayesian and bandit extensions reviewed for the next stage.

1 · Design

Metric and sample size

One decision metric, a minimum detectable effect and error rates fix the sample per arm and the test length.

power analysis
2 · Assignment

Randomize

Users or sessions assigned to arms at random, so time effects hit both arms.

randomization
3 · Analysis

Interval and test

Lift with a confidence interval and a p-value at the planned sample.

z-testCI
4 · Bayesian

Probability to beat

Beta posteriors per arm and the probability that the new variant wins.

Beta-Binomial
5 · Bandit

When traffic is costly

Thompson sampling shifts traffic toward the leader during the test.

Thompson sampling
Stack
LayerTechnologyWhat it does here
DesignSample-size and power formulas for proportions and meansTest length known before launch
AssignmentRandom treatment/control allocationUnbiased comparison through promotions and drift
AnalysisTwo-proportion z-tests, confidence intervalsDecision at the planned sample
ExtensionsBayesian A/B testing, multi-armed bandits (Thompson sampling)Reviewed for platform enhancement
ReportingResult screens for product teamsThe same layout for every experiment
Core formulation
sample per arm   n = ( z_(1−α/2)·√(2p̄(1−p̄)) + z_(1−β)·√(p_A(1−p_A) + p_B(1−p_B)) )² / δ²

test             z = (p̂_B − p̂_A) / √( p̄(1 − p̄)·(1/n_A + 1/n_B) )
interval         (p̂_B − p̂_A) ± z_(1−α/2)·√( p̂_A(1−p̂_A)/n_A + p̂_B(1−p̂_B)/n_B )

Bayesian         p_k ~ Beta(1 + x_k, 1 + n_k − x_k);       P(p_B > p_A) by sampling
Thompson         each round draw θ_k ~ Beta(·) and serve argmax_k θ_k
  • Decide on schedule. The estimate is read at the planned sample, not whenever it first looks significant — which is what keeps the stated error rates true.
  • Two numbers for two questions. The interval says how big the effect is; the posterior probability says how likely the new variant is better.
  • Bandits trade clarity for traffic. Fewer users see the worse arm, at the price of a wider final interval.
In production vs in the live model
ComponentIn productionIn the live model above
PlatformDesign, assignment, analysis and reporting for IPTV and mobile TVOne simulated month of sessions
StatisticsSample size, randomization, tests, Bayesian and bandit methodsThe same formulas, computed in the browser
ConfoundingReal promotions, content drops and driftA planted promotion week the analyst does not know about

Design notes

Decide the sample before the data

The most common failure of informal testing is not the statistics but the stopping: looking every day and calling it when the number first looks good. Fixing the sample size from the minimum detectable effect, and reading the result at that point, is what makes the stated error rates true. The platform shows the running estimate, and still decides on schedule.

Intervals, then probabilities

A confidence interval is the honest summary, and product teams found a second number easier to act on: the posterior probability that B beats A. Reporting both — and explaining that they answer different questions — turned out to matter more for adoption than any single method choice.

When traffic is the scarce thing

For changes where every session shown the worse variant has a cost, multi-armed bandits shift traffic toward the leader while the test runs. They buy fewer lost conversions with a wider interval and a less clean read; the platform treats them as a tool for a specific situation, not a replacement for the test.

Outcome

Controlled experimentsin place of before/after comparisons for IPTV and mobile TV
Standardizeddesign, sample size, randomization and reporting across teams
ExtensibleBayesian testing and bandits reviewed for the next stage

The platform enabled IPTV and mobile TV teams to evaluate service changes using controlled experiments instead of simple before/after comparisons. It standardized experiment quality and strengthened data-driven product decisions.

Limitations

About the demo and confidentiality

Sessions, conversion rates, the promotion and the effect of the variant in the embedded model are simulated. No experiment, metric definition, traffic figure or result from the real platform appears here.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Platform design: metric definition, sample size, randomization, testing and interpretation; review of Bayesian and bandit extensions. Demo simulated for this site.