Media analytics · ForecastingLG Uplus, IPTV movie businessWrite-up October 2026 · 6 min read

Forecasting a film's VOD sales, and the market it sells into

Which films to license, which to make free, when to promote and how much marketing to spend: every one of these decisions in an IPTV movie business rests on expected revenue. This is the three-part forecasting system built for it — one model for a title's total, one for how its sales unfold after release, one for the whole market — and a version you can run.

Built withPythonpandasscikit-learnXGBoostLightGBMk-means · DTWBass diffusionProphet
forecast from day 8 ONE TITLE · DAILY PURCHASES WHOLE MARKET · 26 WEEKS 4 weeks ahead weekend peaks · red = holidays
Left: a title's daily purchases with a Bass diffusion curve fitted to the first week. Right: the market's daily total with its weekly rhythm, holidays and a four-week forecast.

Three questions, three forecasts

IPTV movie decisions — content acquisition, free-movie selection, promotion planning, marketing resource allocation — depend on expected content revenue, and the people making them ask three different questions. Before a licence is signed: how much will this title earn in total? After release: how will that total unfold week by week, so that promotion is timed to the curve rather than against it? And across the catalogue: where is the whole market heading, so that budgets and targets are set against the tide rather than against last month?

A single model answers none of these well. The system was therefore split into three forecasts that share data but not structure.

What was built

Live model, computed in your browser on a generated catalogue of sixty past titles. Left: one new release after its first week, with a forecast that scales the catalogue's average curve and one that combines a feature-based total with a Bass curve fitted to the days seen. Right: the market's daily total with a trend-seasonality-holiday-release model against a seasonal-naive forecast. Switch titles, give the model fourteen days instead of seven, or draw a new catalogue. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

Three forecasts over one title-level data set: a tree-ensemble regressor for each title's total, curve-shape clustering and a diffusion model for its trajectory, and a component model for the whole market.

1 · Data

Titles and the market

VOD purchase history joined with title metadata, theatrical box office and audience ratings collected from public sources.

PythonpandasKOBIS data
2 · Title total

Tree ensembles

XGBoost tuned by grid search with five-fold cross-validation, compared with random forest and LightGBM; permutation importance for the drivers.

XGBoostscikit-learnLightGBM
3 · Curve shape

Cluster, then predict

Eight-week sales-ratio curves clustered with k-means (Euclidean and DTW compared); a multiclass XGBoost predicts a new title's curve type.

k-meansDTWXGBoost
4 · Trajectory

Bass diffusion

Innovation and imitation coefficients fitted to early sales give the post-release curve.

Bass modelnonlinear least squares
5 · Market

Components

Prophet with Korean holidays and a release-event regressor for major titles, ensembled for the total.

Prophet
Stack
LayerTechnologyWhat it does here
DataPython, pandas; public box-office data (KOBIS), audience ratingsTitle-level features joined to VOD sales
Title totalXGBoost (grid search, 5-fold CV), random forest, LightGBM; permutation importanceRevenue over the selling window
Curve typek-means on 8-week sales ratios (Euclidean vs DTW); XGBoost multiclass with cross-entropyWhich demand shape a new title will follow
TrajectoryBass diffusion fitted by nonlinear least squaresWeek-by-week curve after release
MarketProphet with holidays and a release regressor; ensembleTotal daily sales
EvaluationRMSE, weighted MAPE, mBRSE, weighted accuracy; rolling forecast originsErrors weighted toward the titles that matter
Core formulation
title total   log ŷ = f_XGB(box office, admissions, ratings, genre, release timing, …)

curve type    r = 8-week sales ratios;   c = argmin_k ‖r − μ_k‖²           k-means, k = 4 (DTW compared)
              P(c | title) = softmax(f_XGB(title)),    loss = cross-entropy

Bass          F(t) = (1 − e^(−(p+q)t)) / (1 + (q/p)·e^(−(p+q)t)),     sales_t = m·[F(t) − F(t−1)]

market        y(t) = g(t) + s_week(t) + s_year(t) + h_holiday(t) + β·release_t + ε

wMAPE         Σ_i |ŷ_i − y_i| / Σ_i y_i              weights errors by sales volume
  • Weighted metrics. Weighted MAPE and weighted accuracy put more weight on popular titles, which is where acquisition and promotion money goes.
  • Curve types are predictable. Clusters differed in box office, admissions and ratings, so a new title's curve type can be predicted before its sales arrive.
  • Events in the market model. Holidays and blockbuster release dates as regressors were what made the market forecast beat the default component model.
In production vs in the live model
ComponentIn productionIn the live model above
Title totalXGBoost / RF / LightGBM on catalogue featuresLog-linear regression on three generated features
TrajectoryCurve-type clustering + Bass curveBass curve fitted by grid search with the feature total as a prior
MarketProphet with holidays and release events, ensembledThe same additive structure fitted by least squares
DataReal titles, box office and VOD sales60 generated titles and a generated market

Design notes

A total from features, a shape from the first days

Seven days of sales cannot identify a diffusion curve on their own: a fast-decaying blockbuster and a slow-building sleeper can look similar for a week. The feature-based total acts as a prior for the Bass curve's scale, and the early days choose between a front-loaded shape and a word-of-mouth one. As more days arrive, the data takes over from the prior.

Why a parametric curve

A flexible model could fit the catalogue's curves more closely, but the Bass coefficients mean something to the business — a high imitation coefficient is a title that will reward a second promotion push — and a two-parameter curve fitted to a week is far more stable than a free-form one.

Components over black boxes

For the market, an additive decomposition was chosen because planners need to argue with the forecast: how much of next month is trend, how much is the holiday, how much is the release slate. A model that returns a number without those parts gets ignored the first time it is wrong.

Outcome

Threeforecasts — title total, trajectory, market — replacing one rule of thumb
Acquisitionand free-movie selection informed by expected revenue
Promotiontimed to each title's curve rather than the calendar

The project provided a data-driven basis for content acquisition, free-movie selection, promotion planning and marketing resource allocation by forecasting both individual content performance and overall market movement.

Limitations

About the demo and confidentiality

Titles, admissions, sales figures, the holiday calendar and all errors in the embedded model are generated. No content, sales or customer data from any service appears here.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Problem framing, feature design, modelling of all three forecasts. Demo re-implemented on generated data for this site.