Research · Bio AI · Random forest and language modelsSITC 2024 · J ImmunoTher Cancer (abstract)October 2026 · 5 min read

Gut microbiome and immunotherapy response across cancers and cohorts: a random forest against a taxonomy language model

Why a microbiome signal has to be tested cohort by cohort, what a random forest and a language model that reads species as words each bring to it, and why neither of them won everywhere.

Built withNine public metagenomic cohortsEz-Mx platformEzBioCloud databaseSpecies-level taxonomic profilesRandom forestRoBERTa-based language modelThree-fold cross-validationLeave-one-out cross-validationAUROCRandom-effects meta-analysis
ONE PROFILE · SORTED BY ABUNDANCEAUROC BY COHORT · LEAVE-ONE-OUT sp1sp2sp3sp4sp5sp6sp7sp8 species become tokens, most abundant first ROBERTA · PRETRAINED ON 13,418 PROFILES predict DEROSA NSCLCROUTY NSCLCROUTY RCCMCCULLOCH MELSPENCER MELLEE MELFRANKEL MELMATSON MELPETERS MEL 0.30.5 chance0.8 LANGUAGE MODELRANDOM FOREST
For the language model, a gut profile becomes a sentence of species names, most abundant first (left, illustrative). Right: the published leave-one-out AUROC of both models in each of the nine cohorts; the language model is ahead in five, the random forest in three.

Immune checkpoint inhibitors (ICIs) release a brake on the immune system and work against several cancers, yet only some patients benefit. The gut microbiome is a candidate predictor from outside the tumour, but studies in different hospitals and countries have tied response to different bacteria, and the abstract names differences between cohorts as the reason. A signal worth using has to survive a change of cohort, and ideally a change of cancer type.

The study pooled 942 gut metagenomes from patients treated with ICIs for melanoma, non-small cell lung cancer (NSCLC) or renal cell carcinoma (RCC), drawn from nine public cohorts and profiled with the Ez-Mx platform against the EzBioCloud database. Two models predicted response: a random forest, and a RoBERTa-based language model that reads each sample as its species names in order of abundance, pretrained on 13,418 quality-controlled taxonomic profiles and then fine-tuned. Performance was measured as the area under the ROC curve (AUROC) under three-fold and leave-one-out cross-validation, each repeated 100 times.

Leaving out patients with stable disease and profiling at the species level gave the best AUROC. With those settings the language model scored higher than the random forest in five of the nine cohorts, the random forest in three, and one was a tie; Lachnospiraceae species, led by Blautia wexlerae, were among the strongest predictors. I am the third of ten authors, part of the CJ AI Center team that worked with CJ Bioscience researchers on predicting immunotherapy response from the microbiome; the first two authors are from CJ Bioscience, and the abstract does not list individual contributions. This post covers why the evaluation went cohort by cohort, what each model brings, and what a split verdict means.

How it worksThe study at a glance: nine public cohorts across three cancer types, a response label without stable disease, species-level profiles read by a random forest and by a RoBERTa-based language model, AUROC reported cohort by cohort, and the species behind the predictions. All values are from the published abstract.
The study at a glance: nine public cohorts across three cancer types, a response label without stable disease, species-level profiles read by a random forest and by a RoBERTa-based language model, AUROC reported cohort by cohort, and the species behind the predictions. All values are from the published abstract.

Why one cohort is not enough

A gut microbiome profile is a long, sparse list of species whose abundances depend on diet, geography and medication, and on how samples were collected, sequenced and processed. All of that differs between cohorts, and so can the way response was judged. A random split puts patients from every cohort into every fold, so a model can learn those differences and report them as biology.

The study therefore used two schemes. Three-fold cross-validation mixes the samples, so test patients come from familiar cohorts. The leave-one-out results are reported for each named cohort, with a median across cohorts; the abstract does not spell out the unit left out, but that reporting reads as each cohort being scored on its own, the closer test for a biomarker meant for a new hospital. The gap shows: with stable disease excluded, AUROC was 0.653 under three-fold cross-validation and a median of 0.596 under leave-one-out. Repeating each scheme 100 times averages away the luck of any single split.

The label was tested too. Stable disease sits between response and progression, and counting it on either side blurred the signal: under cross-validation the AUROC was 0.583 with stable disease counted as response and 0.637 with it counted as non-response, against 0.653 without it, and the leave-one-out medians followed (0.573 and 0.575 against 0.596). Species-level profiles also beat every other taxonomic rank.

A random forest and a language model that reads species

Inside the methods. Top: a pooled three-fold split shares cohorts between training and test, while scoring each cohort on its own asks the model to work somewhere new; the AUROCs are the study's values for the same task. Bottom: a random forest reads abundances as a table, and the language model reads the same sample as species tokens sorted by abundance (abundances and trees are illustration).
Inside the methods. Top: a pooled three-fold split shares cohorts between training and test, while scoring each cohort on its own asks the model to work somewhere new; the AUROCs are the study's values for the same task. Bottom: a random forest reads abundances as a table, and the language model reads the same sample as species tokens sorted by abundance (abundances and trees are illustration).

A random forest averages the votes of many decision trees, each grown on a resampled set of patients and a random subset of species. That suits a microbiome table: threshold splits make a rule such as "this species above some abundance" cheap to express, the many zeros cause no trouble, and interactions between species are found without being specified. Averaging decorrelated trees keeps the variance down when species far outnumber patients, and the forest ranks its features by contribution, turning a predictor into a list of candidates.

The language model reads the same sample as a sentence whose tokens are species names, sorted from most to least abundant. RoBERTa is a Transformer encoder, and its self-attention represents each species in the context of the others. Pretraining on 13,418 profiles, far more than the 942 labelled samples, lets it learn which species tend to occur together before it sees a response label; fine-tuning then adapts it to the response task. It gives up exact abundances, keeps their order, and gains community structure learned from unlabelled data, which a forest sees only through its splits.

Results

Cohort, leave-one-outCancerLanguage model AUROCRandom forest AUROC
Derosa 2022NSCLC0.6390.611
Routy 2018NSCLC0.6760.660
Routy 2018RCC0.6580.560
McCulloch 2022Melanoma0.6180.564
Spencer 2021Melanoma0.6040.554
Lee 2022Melanoma0.5760.580
Frankel 2017Melanoma0.5510.635
Matson 2018Melanoma0.3050.384
Peters 2019Melanoma0.7440.801

The language model led in all three non-melanoma cohorts, the random forest in three melanoma cohorts. In Matson 2018 both scored below 0.5: patterns learned elsewhere ranked that cohort's patients worse than chance. Lachnospiraceae species stood out among the predictive features; in a random-effects meta-analysis, which pools cohorts while letting the true effect differ between them, Blautia wexlerae (p = 0.009), Anaerostipes hadrus (p = 0.053) and Fusicatenibacter saccharivorans (p = 0.016) were enriched in responders. The authors read the AUROCs as reasonable for a hard problem and the species as hints of a role in immune modulation.

What I learned

The evaluation scheme is part of the result. The same task scored 0.653 with cohorts shared between training and test and a median of 0.596 with each cohort scored on its own. For a biomarker meant for a new hospital the second number is the honest one, and per cohort it shows where a model breaks.

Label definitions can matter as much as model choice. Counting stable disease as response cost 0.07 of AUROC under cross-validation, more than the gap between the two models in six of the nine cohorts.

Two models that fail in different places are worth combining. The abstract reports no ensemble, but its split verdict is the classic argument for one: a tree model on abundances and a sequence model on species order make different mistakes, and averaging them is usually steadier than betting on either. The wider CJ work behind the poster is described in the immunotherapy-response project write-up.

Limitations

All numbers come from the published SITC 2024 abstract in the Journal for ImmunoTherapy of Cancer. The figures are drawn for this site, values marked as illustration are invented to explain the methods, and no patient-level data are shown or shared.

Taehee Lee · Third author · CJ AI Center teamThird of ten authors, as a member of the CJ AI Center team that worked with CJ Bioscience researchers on predicting immunotherapy response from microbiome data. Published as a conference abstract in Journal for ImmunoTherapy of Cancer 12(Suppl 2): A1421, 2024 (SITC 2024, poster 1265). doi:10.1136/jitc-2024-sitc2024.1265