Live model · simulated month of sessions · computed in your browser

Did the new home screen work? Two ways to answer, one of them right

A streaming home screen gets a new recommendation shelf. The quickest answer is to launch it and compare the two weeks after with the two weeks before — but promotions, content drops and drift arrive in the same weeks. The platform instead plans the sample, randomizes every session, and decides at the planned point with a confidence interval, a Bayesian probability and, if traffic is precious, a bandit.

Experimentation platform

Simulating the month…

Before · launch to everyone, compare before vs after
After · randomized test with a planned sample
old screen (A)new screen (B)content promotion, unknown to the analyst95% confidence bandtrue effect
1 · Plan

Metric, minimum effect, sample size

The team names the metric (play-starts per session), the smallest lift worth shipping (0.4 pp on an 8% base) and the error rates (5% false positive, 80% power). That fixes the sample per arm — and therefore how long the test runs — before anyone sees a number.

2 · Randomize and wait

Both arms live through the same weeks

Every session is assigned to A or B at random, so promotions, weekends and drift hit both arms equally and cancel in the difference. The estimate is read at the planned sample, not whenever it first looks good.

3 · Decide, with options

Interval, probability, bandit

The result is a lift with a confidence interval and a p-value, plus a Bayesian probability that B beats A, which product teams find easier to act on. Where exposing users to a worse variant is costly, Thompson sampling shifts traffic toward the leader during the test — at the price of a wider interval.