Finding the conditions that actually move yield
A fermentation batch records hundreds of variables, and many of them move with yield without causing it. This is how causal discovery and an in-silico simulator turned one-off analyses into a repeatable way to find operating conditions worth trying on a real line — and a small version you can run.
Correlated is not controllable
Bio-processes have many variables with complex interactions. Data was spread across systems, and many variables were managed by hand, so building an analyzable dataset was itself the first step. The existing practice relied on expert experience and one-off analyses — and because process conditions drive manufacturing cost directly, the conditions had to be explainable and reliable enough to test on an actual production line.
Correlation and simple tests are a poor guide here. A variable that shares a cause with yield passes every screen and changes nothing when you adjust it. A condition with an optimum in the middle of its range — temperature is the classic case — looks unimportant to a straight-line test. The goal was not "which variables move with yield" but "which variables, if we set them, would move yield".
From data mart to do-operator
- Understand the process with its experts — key variables, stages, how each measurement is produced, and which conditions can actually be controlled.
- Build the data mart — integrate system and hand-collected sources into one batch-level table, with derived features such as concentrations and feed rates per phase.
- Learn cause-and-effect structure — causal discovery with the process order as prior knowledge: later stages cannot cause earlier ones, so a large share of possible links is ruled out before learning, and expert-supplied relationships can be forced in or out.
- Simulate interventions — with the structure and fitted structural equations, set a variable to a value and propagate through its descendants to estimate the effect on yield.
Live model, computed on a Python server from 450 generated batches with a known structure planted inside. Left: a correlation screen. Right: the structure learned under the process-order prior. Below: the effect of setting a variable — the causal model's curve, the planted truth, and what the correlation line would have predicted. Click any controllable variable; "New batches" draws a fresh dataset. Open the live model on its own page ↗
Architecture, stack and core formulation
From long-format batch records to a causal graph and a simulator process experts can query: screening, prior knowledge as a constraint matrix, deep end-to-end causal discovery, and do-interventions.
One row per batch
Time-stamped measurements pivoted per batch and joined with media data; derived concentrations, feed rates and a yield-per-time target.
Shortlist candidates
Missing-value and outlier handling, Welch t-tests between top and bottom batches, quadratic single-variable tests and random-forest SHAP.
Process order as constraints
Experts arrange variables by process stage on a drag-and-drop screen; edges against process order are forbidden.
Graph and equations together
DECI learns the causal graph and nonlinear structural equations jointly, under the constraint matrix.
do-interventions
Sample yield under do(X = x) across a grid and compare two settings with an effect size and a test.
| Layer | Technology | What it does here |
|---|---|---|
| Data | pandas, openpyxl | Batch × time-point records pivoted to batch features |
| Screening | SciPy (Welch t), statsmodels (OLS F-test), random forest + SHAP | From hundreds of columns to a candidate set |
| Constraints | Expert ordering → {learn, forbid, force} matrix | Removes edges that contradict process order |
| Causal discovery | causica DECI on PyTorch / Lightning | Graph + nonlinear structural equations with an acyclicity penalty |
| Simulator | Streamlit multi-screen app, interactive graph view, plotly | Process experts test hypotheses before touching the line |
constraint matrix C ∈ {learn, 0 = forbid, 1 = force}^((d+1)×(d+1))
C[i,j] = 0 if stage(i) ≥ stage(j) no edges within or against process order
DECI x_j = f_j(x_pa(j); θ) + ε_j, G ~ q_φ(G)
max E_q[ log p(x | G, θ) ] − λ·‖G‖₁ sparsity
s.t. h(G) = tr(e^(G∘G)) − d = 0 acyclicity, augmented Lagrangian
effect curve E[ yield | do(X = x) ], x = p10 … p90
two-point test ATE = E[Y | do(X = b)] − E[Y | do(X = a)]- Prior knowledge does most of the pruning. Forbidding edges against process order removes a large share of candidate edges before learning starts.
- Screening is a filter, not an answer. A variable can pass correlation and t-tests and still have no causal path to yield — the live model shows exactly that case.
- Simulator for the experts. Effect curves and two-point comparisons let process engineers check a hypothesis before running a batch.
| Component | In production | In the live model above |
|---|---|---|
| Discovery | DECI (deep end-to-end causal inference) under process-order constraints | Order-constrained forward selection by BIC with linear + quadratic terms (NumPy) |
| Screening | t-tests, quadratic F-tests, random forest + SHAP | Correlation and Welch t screen |
| Simulator | Streamlit app over the learned model | do-intervention curves from the learned equations |
| Data | Real batch records from three plants | A generated process with a planted causal graph |
The in-silico simulator
The analysis was delivered as a tool, not a report. Process experts could load data, choose variables, encode the stage order and their own known relationships, train the causal model, and then ask "what if" questions: the effect curve of a condition across its realistic range, or a direct comparison between a control value and a treatment value with its estimated effect and uncertainty. Hypotheses that survived the simulator were the ones taken to the line.
In production the structure was learned with a deep structural-equation method that fits the graph and nonlinear equations together under an acyclicity constraint. The live model above keeps the same three ideas — order constraints, nonlinear equations, do-interventions — with a lighter learner so it can run in a second on a small server.
What it delivered
The system was applied at three global sites. It produced hypotheses that could be tested on real production lines, and some of the conditions it found were adopted as standard process conditions. It also turned a one-off analysis into a repeatable improvement system, and it was connected to a technology-licensing agreement.
A further outcome was an agenda for manufacturing data itself: digitizing hand-managed records, improving data management, and building automated pipelines and a process knowledge base — so the next round of analysis starts from better data.
Limitations
- Causal discovery from observational batches depends on the prior and on the variables measured; an unmeasured common cause can still mislead it. Conditions were validated on the line before adoption.
- The simulator estimates average effects within the observed operating range; it does not extrapolate safely beyond it.
- The live model's planted structure, variable names and effect sizes are invented; its accuracy on synthetic data says nothing about accuracy on a real process.
About the demo and confidentiality
Every batch, variable, unit and coefficient in the embedded model is generated. No product, strain, media composition, process condition, plant data or result from the real project appears here.