Manufacturing · Causal inference · SimulationCJ AI Center, with CJ CheilJedangWrite-up October 2026 · 8 min read

Finding the conditions that actually move yield

A fermentation batch records hundreds of variables, and many of them move with yield without causing it. This is how causal discovery and an in-silico simulator turned one-off analyses into a repeatable way to find operating conditions worth trying on a real line — and a small version you can run.

Built withPythonpandasstatsmodelsSHAPcausica · DECIPyTorchStreamlitplotly
FLASKSEEDMEDIAMAIN · EARLYMAIN · MIDYIELD
Stages left to right; arrows only point forward. The dashed box correlates strongly with yield and causes none of it. The orange curve is an effect with an optimum in the middle — invisible to a straight-line screen.

Correlated is not controllable

Bio-processes have many variables with complex interactions. Data was spread across systems, and many variables were managed by hand, so building an analyzable dataset was itself the first step. The existing practice relied on expert experience and one-off analyses — and because process conditions drive manufacturing cost directly, the conditions had to be explainable and reliable enough to test on an actual production line.

Correlation and simple tests are a poor guide here. A variable that shares a cause with yield passes every screen and changes nothing when you adjust it. A condition with an optimum in the middle of its range — temperature is the classic case — looks unimportant to a straight-line test. The goal was not "which variables move with yield" but "which variables, if we set them, would move yield".

From data mart to do-operator

  1. Understand the process with its experts — key variables, stages, how each measurement is produced, and which conditions can actually be controlled.
  2. Build the data mart — integrate system and hand-collected sources into one batch-level table, with derived features such as concentrations and feed rates per phase.
  3. Learn cause-and-effect structure — causal discovery with the process order as prior knowledge: later stages cannot cause earlier ones, so a large share of possible links is ruled out before learning, and expert-supplied relationships can be forced in or out.
  4. Simulate interventions — with the structure and fitted structural equations, set a variable to a value and propagate through its descendants to estimate the effect on yield.

Live model, computed on a Python server from 450 generated batches with a known structure planted inside. Left: a correlation screen. Right: the structure learned under the process-order prior. Below: the effect of setting a variable — the causal model's curve, the planted truth, and what the correlation line would have predicted. Click any controllable variable; "New batches" draws a fresh dataset. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

From long-format batch records to a causal graph and a simulator process experts can query: screening, prior knowledge as a constraint matrix, deep end-to-end causal discovery, and do-interventions.

1 · Batch table

One row per batch

Time-stamped measurements pivoted per batch and joined with media data; derived concentrations, feed rates and a yield-per-time target.

pandasJupyter
2 · Screening

Shortlist candidates

Missing-value and outlier handling, Welch t-tests between top and bottom batches, quadratic single-variable tests and random-forest SHAP.

SciPystatsmodelsSHAP
3 · Prior knowledge

Process order as constraints

Experts arrange variables by process stage on a drag-and-drop screen; edges against process order are forbidden.

Streamlitconstraint matrix
4 · Discovery

Graph and equations together

DECI learns the causal graph and nonlinear structural equations jointly, under the constraint matrix.

causica DECIPyTorch
5 · Simulator

do-interventions

Sample yield under do(X = x) across a grid and compare two settings with an effect size and a test.

Streamlitplotly
Stack
LayerTechnologyWhat it does here
Datapandas, openpyxlBatch × time-point records pivoted to batch features
ScreeningSciPy (Welch t), statsmodels (OLS F-test), random forest + SHAPFrom hundreds of columns to a candidate set
ConstraintsExpert ordering → {learn, forbid, force} matrixRemoves edges that contradict process order
Causal discoverycausica DECI on PyTorch / LightningGraph + nonlinear structural equations with an acyclicity penalty
SimulatorStreamlit multi-screen app, interactive graph view, plotlyProcess experts test hypotheses before touching the line
Core formulation
constraint matrix   C ∈ {learn, 0 = forbid, 1 = force}^((d+1)×(d+1))
                    C[i,j] = 0   if stage(i) ≥ stage(j)        no edges within or against process order

DECI   x_j = f_j(x_pa(j); θ) + ε_j,      G ~ q_φ(G)
       max  E_q[ log p(x | G, θ) ]  −  λ·‖G‖₁                    sparsity
       s.t. h(G) = tr(e^(G∘G)) − d = 0                          acyclicity, augmented Lagrangian

effect curve     E[ yield | do(X = x) ],    x = p10 … p90
two-point test   ATE = E[Y | do(X = b)] − E[Y | do(X = a)]
  • Prior knowledge does most of the pruning. Forbidding edges against process order removes a large share of candidate edges before learning starts.
  • Screening is a filter, not an answer. A variable can pass correlation and t-tests and still have no causal path to yield — the live model shows exactly that case.
  • Simulator for the experts. Effect curves and two-point comparisons let process engineers check a hypothesis before running a batch.
In production vs in the live model
ComponentIn productionIn the live model above
DiscoveryDECI (deep end-to-end causal inference) under process-order constraintsOrder-constrained forward selection by BIC with linear + quadratic terms (NumPy)
Screeningt-tests, quadratic F-tests, random forest + SHAPCorrelation and Welch t screen
SimulatorStreamlit app over the learned modeldo-intervention curves from the learned equations
DataReal batch records from three plantsA generated process with a planted causal graph

The in-silico simulator

The analysis was delivered as a tool, not a report. Process experts could load data, choose variables, encode the stage order and their own known relationships, train the causal model, and then ask "what if" questions: the effect curve of a condition across its realistic range, or a direct comparison between a control value and a treatment value with its estimated effect and uncertainty. Hypotheses that survived the simulator were the ones taken to the line.

In production the structure was learned with a deep structural-equation method that fits the graph and nonlinear equations together under an acyclicity constraint. The live model above keeps the same three ideas — order constraints, nonlinear equations, do-interventions — with a lighter learner so it can run in a second on a small server.

What it delivered

≈ $1.86Mconfirmed cost savings across three global sites
$773Kof it realized through the in-silico simulator
5patents on causal condition discovery and process optimization

The system was applied at three global sites. It produced hypotheses that could be tested on real production lines, and some of the conditions it found were adopted as standard process conditions. It also turned a one-off analysis into a repeatable improvement system, and it was connected to a technology-licensing agreement.

A further outcome was an agenda for manufacturing data itself: digitizing hand-managed records, improving data management, and building automated pipelines and a process knowledge base — so the next round of analysis starts from better data.

Limitations

About the demo and confidentiality

Every batch, variable, unit and coefficient in the embedded model is generated. No product, strain, media composition, process condition, plant data or result from the real project appears here.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterProblem framing with process experts, data mart, causal discovery and inference, in-silico simulator, patents. Demo re-implemented on synthetic data for this site.