Forecasting a film's VOD sales, and the market it sells into
Which films to license, which to make free, when to promote and how much marketing to spend: every one of these decisions in an IPTV movie business rests on expected revenue. This is the three-part forecasting system built for it — one model for a title's total, one for how its sales unfold after release, one for the whole market — and a version you can run.
Three questions, three forecasts
IPTV movie decisions — content acquisition, free-movie selection, promotion planning, marketing resource allocation — depend on expected content revenue, and the people making them ask three different questions. Before a licence is signed: how much will this title earn in total? After release: how will that total unfold week by week, so that promotion is timed to the curve rather than against it? And across the catalogue: where is the whole market heading, so that budgets and targets are set against the tide rather than against last month?
A single model answers none of these well. The system was therefore split into three forecasts that share data but not structure.
What was built
- Title total. Tree-based models — Random Forest, XGBoost, LightGBM — on title features such as theatrical performance, release timing and genre, trained on the back catalogue, predict a new title's revenue over its selling window.
- Post-release trajectory. The Bass diffusion model describes how a new product is adopted by a population through two interpretable coefficients: innovation (people who buy on their own initiative) and imitation (people who follow others). Fitted to the first days of sales, it gives the shape of the remaining curve.
- Total market. Prophet decomposes the sum of all titles' revenue into trend, weekly and yearly seasonality and holiday effects, with external regressors for the release calendar.
Live model, computed in your browser on a generated catalogue of sixty past titles. Left: one new release after its first week, with a forecast that scales the catalogue's average curve and one that combines a feature-based total with a Bass curve fitted to the days seen. Right: the market's daily total with a trend-seasonality-holiday-release model against a seasonal-naive forecast. Switch titles, give the model fourteen days instead of seven, or draw a new catalogue. Open the live model on its own page ↗
Architecture, stack and core formulation
Three forecasts over one title-level data set: a tree-ensemble regressor for each title's total, curve-shape clustering and a diffusion model for its trajectory, and a component model for the whole market.
Titles and the market
VOD purchase history joined with title metadata, theatrical box office and audience ratings collected from public sources.
Tree ensembles
XGBoost tuned by grid search with five-fold cross-validation, compared with random forest and LightGBM; permutation importance for the drivers.
Cluster, then predict
Eight-week sales-ratio curves clustered with k-means (Euclidean and DTW compared); a multiclass XGBoost predicts a new title's curve type.
Bass diffusion
Innovation and imitation coefficients fitted to early sales give the post-release curve.
Components
Prophet with Korean holidays and a release-event regressor for major titles, ensembled for the total.
| Layer | Technology | What it does here |
|---|---|---|
| Data | Python, pandas; public box-office data (KOBIS), audience ratings | Title-level features joined to VOD sales |
| Title total | XGBoost (grid search, 5-fold CV), random forest, LightGBM; permutation importance | Revenue over the selling window |
| Curve type | k-means on 8-week sales ratios (Euclidean vs DTW); XGBoost multiclass with cross-entropy | Which demand shape a new title will follow |
| Trajectory | Bass diffusion fitted by nonlinear least squares | Week-by-week curve after release |
| Market | Prophet with holidays and a release regressor; ensemble | Total daily sales |
| Evaluation | RMSE, weighted MAPE, mBRSE, weighted accuracy; rolling forecast origins | Errors weighted toward the titles that matter |
title total log ŷ = f_XGB(box office, admissions, ratings, genre, release timing, …)
curve type r = 8-week sales ratios; c = argmin_k ‖r − μ_k‖² k-means, k = 4 (DTW compared)
P(c | title) = softmax(f_XGB(title)), loss = cross-entropy
Bass F(t) = (1 − e^(−(p+q)t)) / (1 + (q/p)·e^(−(p+q)t)), sales_t = m·[F(t) − F(t−1)]
market y(t) = g(t) + s_week(t) + s_year(t) + h_holiday(t) + β·release_t + ε
wMAPE Σ_i |ŷ_i − y_i| / Σ_i y_i weights errors by sales volume- Weighted metrics. Weighted MAPE and weighted accuracy put more weight on popular titles, which is where acquisition and promotion money goes.
- Curve types are predictable. Clusters differed in box office, admissions and ratings, so a new title's curve type can be predicted before its sales arrive.
- Events in the market model. Holidays and blockbuster release dates as regressors were what made the market forecast beat the default component model.
| Component | In production | In the live model above |
|---|---|---|
| Title total | XGBoost / RF / LightGBM on catalogue features | Log-linear regression on three generated features |
| Trajectory | Curve-type clustering + Bass curve | Bass curve fitted by grid search with the feature total as a prior |
| Market | Prophet with holidays and release events, ensembled | The same additive structure fitted by least squares |
| Data | Real titles, box office and VOD sales | 60 generated titles and a generated market |
Design notes
A total from features, a shape from the first days
Seven days of sales cannot identify a diffusion curve on their own: a fast-decaying blockbuster and a slow-building sleeper can look similar for a week. The feature-based total acts as a prior for the Bass curve's scale, and the early days choose between a front-loaded shape and a word-of-mouth one. As more days arrive, the data takes over from the prior.
Why a parametric curve
A flexible model could fit the catalogue's curves more closely, but the Bass coefficients mean something to the business — a high imitation coefficient is a title that will reward a second promotion push — and a two-parameter curve fitted to a week is far more stable than a free-form one.
Components over black boxes
For the market, an additive decomposition was chosen because planners need to argue with the forecast: how much of next month is trend, how much is the holiday, how much is the release slate. A model that returns a number without those parts gets ignored the first time it is wrong.
Outcome
The project provided a data-driven basis for content acquisition, free-movie selection, promotion planning and marketing resource allocation by forecasting both individual content performance and overall market movement.
Limitations
- Online sales are volatile and sensitive to external events; the forecasts are decision support with error bands, not point predictions to be held to.
- The Bass model assumes one adoption process; titles with a second life (an award, a sequel's release) need a re-fit.
- The live model's catalogue is generated with the same mechanics the forecasts assume, which flatters them; real catalogues are messier, and the production title-total model used tree ensembles rather than the regression shown here.
About the demo and confidentiality
Titles, admissions, sales figures, the holiday calendar and all errors in the embedded model are generated. No content, sales or customer data from any service appears here.