A streaming home screen gets a new recommendation shelf. The quickest answer is to launch it and compare the two weeks after with the two weeks before — but promotions, content drops and drift arrive in the same weeks. The platform instead plans the sample, randomizes every session, and decides at the planned point with a confidence interval, a Bayesian probability and, if traffic is precious, a bandit.
The team names the metric (play-starts per session), the smallest lift worth shipping (0.4 pp on an 8% base) and the error rates (5% false positive, 80% power). That fixes the sample per arm — and therefore how long the test runs — before anyone sees a number.
Every session is assigned to A or B at random, so promotions, weekends and drift hit both arms equally and cancel in the difference. The estimate is read at the planned sample, not whenever it first looks good.
The result is a lift with a confidence interval and a p-value, plus a Bayesian probability that B beats A, which product teams find easier to act on. Where exposing users to a worse variant is costly, Thompson sampling shifts traffic toward the leader during the test — at the price of a wider interval.