From a Christmas-cake forecast to every store, every day
The request was to forecast Christmas cakes. The decision stores actually face is how much of each category to bake every day of the year. This is the project that re-scoped one into the other — store-level, category-level, year-round forecasts and production recommendations for more than 300 bakeries — and measured a 5.7% revenue lift in testing.
Re-scoping the question
Tous les Jours stores have different sales patterns depending on location, customer base, day-of-week effects, seasonality, promotions, weather and events. The initial discussion focused on forecasting a Christmas-season item. A seasonal or single-item forecast, however good, does not help a store decide what to bake on an ordinary Tuesday — and those ordinary days add up to most of the year's revenue.
I redefined the task as store-level, category-level, year-round demand forecasting and production recommendation. The target expanded from a single Christmas cake to mid-level product categories covering most revenue, and from a limited number of directly operated stores to more than 300 stores.
What was built
- Data that stores actually generate. Required source data was identified and pipelines built for store sales mix, trends, customer characteristics, payment methods, discounts and external commercial-district data. I developed OCR code and contributed to collecting commercial-district analysis data for all stores through CJ Foodville's RPA.
- Forecasting models compared. Time-series models including Temporal Fusion Transformer, PatchTST and DeepAR were compared for store × category forecasts.
- A dashboard for the people who decide. A Streamlit dashboard shows store characteristics, sales rankings, sales mix, payment and discount information, and store- and category-level forecasts.
- Automation end to end. Data collection, forecasting, dashboard generation, result delivery and evaluation run as one automated loop.
Live model, computed in your browser on 300 generated stores in four district types and four categories, with a year of history including last Christmas. Left: baking last week's same-day sales plus 10%. Right: one log-linear model across all stores and categories, with an 80% interval and production at each category's profit-maximizing quantile. Switch between an ordinary fortnight and Christmas, pick a category, or look at another store; chain-wide numbers cover every store, category and day of the two test weeks. The model and the production rule are stand-ins for the project's, chosen to run in a browser. Open the live model on its own page ↗
Architecture, stack and core formulation
An automated loop from data collection to store-level recommendations: pipelines and OCR/RPA for the data, deep time-series models for store × category forecasts, and a dashboard for the stores.
Sales and context
Sales mix, trends, customers, payment methods, discounts and commercial-district data for every store, gathered with OCR code and RPA.
Store features
Store characteristics and clustering; series at store × mid-level category × day.
Deep time-series models
Temporal Fusion Transformer, PatchTST and DeepAR compared for year-round forecasts.
How much to bake
Forecasts turned into production recommendations for each store and category.
Dashboard and automation
A Streamlit dashboard; collection, forecasting, delivery and evaluation automated end to end.
| Layer | Technology | What it does here |
|---|---|---|
| Data collection | Data pipelines, OCR code, RPA for commercial-district data | Every store's data, refreshed automatically |
| Features | Store clustering, sales mix, payment and discount features | What makes stores alike or different |
| Forecasting | Temporal Fusion Transformer, PatchTST, DeepAR (compared) | Store × category daily forecasts with intervals |
| Dashboard | Streamlit | Store characteristics, rankings, sales mix, forecasts |
| Operations | Automated collection → forecast → delivery → evaluation | An always-on system, licensed as technology |
target ŷ_(s,c,t+1 … t+14) = f_θ( history_(s,c), store features, calendar, promotions ) TFT variable selection + LSTM encoder–decoder + interpretable attention; quantile loss PatchTST series → patches of length P → Transformer encoder (channel-independent) DeepAR y_t ~ p(· | μ_t, σ_t), (μ_t, σ_t) = g(h_t), h_t = RNN(h_(t−1), y_(t−1), x_t) quantity q* = F̂⁻¹( (price − cost) / price ) newsvendor rule in the live model
- The re-scoping was the design. From one Christmas item to mid-level categories covering most revenue, and from a few stores to 300+.
- Probabilistic models for a decision. Quantile and distributional forecasts give the interval a production rule needs, not just a point.
- Automation is part of the model. Without automated collection and delivery a daily forecast is a report; with it, it is an operating system for stores.
| Component | In production | In the live model above |
|---|---|---|
| Models | TFT, PatchTST, DeepAR compared | One regularized log-linear model across all stores |
| Data | 300+ stores: sales, payments, discounts, commercial district | 300 generated stores, four categories |
| Recommendation | Production recommendations via the dashboard | Newsvendor quantile per category |
| Result | +5.7% revenue in testing | Computed live against a rule of thumb |
Design notes
One model across stores
Three hundred stores with a handful of categories each is too many series to model one at a time and too few days per series to learn rare events from. A global model learns the shared structure — what an office district does on a Saturday, what a holiday does to a residential street — once, and lets each store-category contribute only its own level. The live model does the same with a regularized log-linear regression.
From forecast to quantity
A forecast is not a production plan. Stores need a number, and the right number depends on what is worse: an unsold loaf or a customer turned away. The demo makes that explicit with a quantile rule; in the project, forecasts were delivered to stores with intervals and recommendations through the dashboard.
Always-on beats once-a-month
The revenue improvement matters because it is continuous. Large event-driven campaigns run perhaps twice a month; a better daily production decision applies every day, in every store.
Outcome
The key contribution was shifting the problem from a seasonal item forecast to a regular demand-forecasting and production-recommendation system that can be used in store operations, with the whole loop — data collection, forecasting, dashboard, delivery and evaluation — automated.
Limitations
- Sales are censored by what was baked: a sold-out day hides how much more could have sold, and forecasts trained on sales inherit that bias unless it is modelled.
- New stores and new products have no history; they borrow from similar stores, which works for districts and less well for genuinely new items.
- The live model's stores, sales and prices are generated, and its numbers come from a simpler model than the transformer-based ones compared in the project; in its ordinary fortnight the rule of thumb is slightly better on two-week totals, which the demo shows rather than hides.
About the demo and confidentiality
Stores, districts, categories, prices, weather and sales in the embedded model are generated. No store, sales, customer or commercial-district data from Tous les Jours or CJ Foodville appears here.