An experimentation platform for IPTV and mobile TV
New features, screens, recommendation logic and marketing messages were being judged by comparing the weeks after a launch with the weeks before — a method that cannot tell a good change from a good week. This is the platform designed to replace it: a repeatable way to plan, randomize, test and read experiments, built for product teams rather than statisticians.
Good change, or good week?
IPTV and mobile TV services needed a repeatable experimentation platform to evaluate new features, screens, recommendation logic and marketing messages. Without standardized experimental design, results could be biased or confused with natural fluctuation: a new home screen launched in the same fortnight as a content promotion gets credit for the promotion, and a change launched into a quiet week gets blamed for the quiet. Teams were making real decisions on comparisons that could not, in principle, answer the question.
What was built
The platform covers the whole life of an experiment, in the order a product team actually meets it:
- Objective and metric definition — what the change is supposed to move, and the one metric the decision will be read from.
- Sample-size calculation — from the baseline, the smallest effect worth shipping and the agreed error rates, so the test's length is known before it starts.
- Treatment and control randomization — every user or session assigned at random, so that whatever else happens that month happens to both arms.
- Statistical testing and result interpretation — the lift with its confidence interval and p-value, presented so that a non-statistician reads the right conclusion.
- Next-stage methods — Bayesian testing and multi-armed bandit approaches were reviewed for future enhancement, for teams that want a probability rather than a p-value and for cases where exposing users to a worse variant is costly.
Live simulation in your browser. One month of sessions with a weekend lift, drift and a content promotion the analyst does not know about. Left: the new screen launched to everyone on day 15 and judged before versus after. Right: the platform — planned sample, daily randomization, the cumulative lift with its 95% band, a decision at the planned point, the Bayesian probability that B wins, and what Thompson sampling would have done with the same traffic. Set B's true effect and draw a new month. Open the live model on its own page ↗
Architecture, stack and core formulation
The statistical core of an experimentation platform for IPTV and mobile TV: design, assignment, analysis and reporting, with Bayesian and bandit extensions reviewed for the next stage.
Metric and sample size
One decision metric, a minimum detectable effect and error rates fix the sample per arm and the test length.
Randomize
Users or sessions assigned to arms at random, so time effects hit both arms.
Interval and test
Lift with a confidence interval and a p-value at the planned sample.
Probability to beat
Beta posteriors per arm and the probability that the new variant wins.
When traffic is costly
Thompson sampling shifts traffic toward the leader during the test.
| Layer | Technology | What it does here |
|---|---|---|
| Design | Sample-size and power formulas for proportions and means | Test length known before launch |
| Assignment | Random treatment/control allocation | Unbiased comparison through promotions and drift |
| Analysis | Two-proportion z-tests, confidence intervals | Decision at the planned sample |
| Extensions | Bayesian A/B testing, multi-armed bandits (Thompson sampling) | Reviewed for platform enhancement |
| Reporting | Result screens for product teams | The same layout for every experiment |
sample per arm n = ( z_(1−α/2)·√(2p̄(1−p̄)) + z_(1−β)·√(p_A(1−p_A) + p_B(1−p_B)) )² / δ² test z = (p̂_B − p̂_A) / √( p̄(1 − p̄)·(1/n_A + 1/n_B) ) interval (p̂_B − p̂_A) ± z_(1−α/2)·√( p̂_A(1−p̂_A)/n_A + p̂_B(1−p̂_B)/n_B ) Bayesian p_k ~ Beta(1 + x_k, 1 + n_k − x_k); P(p_B > p_A) by sampling Thompson each round draw θ_k ~ Beta(·) and serve argmax_k θ_k
- Decide on schedule. The estimate is read at the planned sample, not whenever it first looks significant — which is what keeps the stated error rates true.
- Two numbers for two questions. The interval says how big the effect is; the posterior probability says how likely the new variant is better.
- Bandits trade clarity for traffic. Fewer users see the worse arm, at the price of a wider final interval.
| Component | In production | In the live model above |
|---|---|---|
| Platform | Design, assignment, analysis and reporting for IPTV and mobile TV | One simulated month of sessions |
| Statistics | Sample size, randomization, tests, Bayesian and bandit methods | The same formulas, computed in the browser |
| Confounding | Real promotions, content drops and drift | A planted promotion week the analyst does not know about |
Design notes
Decide the sample before the data
The most common failure of informal testing is not the statistics but the stopping: looking every day and calling it when the number first looks good. Fixing the sample size from the minimum detectable effect, and reading the result at that point, is what makes the stated error rates true. The platform shows the running estimate, and still decides on schedule.
Intervals, then probabilities
A confidence interval is the honest summary, and product teams found a second number easier to act on: the posterior probability that B beats A. Reporting both — and explaining that they answer different questions — turned out to matter more for adoption than any single method choice.
When traffic is the scarce thing
For changes where every session shown the worse variant has a cost, multi-armed bandits shift traffic toward the leader while the test runs. They buy fewer lost conversions with a wider interval and a less clean read; the platform treats them as a tool for a specific situation, not a replacement for the test.
Outcome
The platform enabled IPTV and mobile TV teams to evaluate service changes using controlled experiments instead of simple before/after comparisons. It standardized experiment quality and strengthened data-driven product decisions.
Limitations
- A randomized test answers the question it was designed for; interference between arms, novelty effects and metrics that move slowly still need judgment.
- Power is a promise about the long run: with 80% power, one month in five will miss a real effect of the minimum size, as the live model shows when you draw new months.
- The embedded simulation uses a single conversion metric and a normal approximation to the posterior; the platform handled several metric types and used exact methods where sample sizes required them.
About the demo and confidentiality
Sessions, conversion rates, the promotion and the effect of the variant in the embedded model are simulated. No experiment, metric definition, traffic figure or result from the real platform appears here.