Microbiome signals of immunotherapy response, and a search engine for proteins
Immune checkpoint inhibitors work for some patients and not others, and the gut microbiome is one of the most promising places to look for the difference. This is the work with CJ Bioscience researchers that turned more than 100 million protein sequences into a searchable embedding space and used it to predict response and surface biomarker candidates — presented at SITC 2024.
Why the microbiome, and why it is hard
Immune checkpoint inhibitors do not work equally for all patients, and response rates are limited. Predicting response before treatment, and identifying the biomarkers related to it, can support clinical decision-making and drug development. For CJ Bioscience's biomarker platform strategy, that meant AI-based analysis to predict immunotherapy response from microbiome data and to explore biomarker candidates.
Two things make this difficult. Microbiome signatures found in one cohort often fail in another, because populations carry different species. And the underlying data — protein sequences, metagenomes, microbiome profiles — is enormous, so any representation of what the microbes can do must be built and searched at scale.
What was built
- Methods surveyed with the researchers. Working closely with CJ Bioscience researchers, I evaluated bioinformatics deep-learning methods: ESM2 and gLM protein embeddings, AlphaFold and ESMFold structure prediction, and MMseqs2 and Foldseek clustering.
- A representation at scale. More than 100 million protein sequences were clustered into about 6 million clusters, embedded, and validated for sequence classification, function inference and search. Embedding processing time fell by about 60× against the baseline.
- A searchable store. A vector-database structure searches about 6 million sequences within 1 to 2 seconds.
- Response prediction and candidates. A microbiome-based response model, and a structure for surfacing biomarker candidates rather than only scores. The related SITC 2024 study used 942 samples from nine public metagenomic cohorts, comparing random forests, a microbiome taxonomic language model, BMCARP and ensembles, and combined species-level profiles with functional contig embeddings.
Live model, computed in your browser on generated metagenomes the size of the public data used in the SITC poster: 942 samples in nine cohorts, each with its own species mix, and a response that depends on a function carried by different species in different cohorts. Top: a species-only model and a model with protein-family clusters, each trained on eight cohorts and tested on the ninth. Bottom: the same clusters used as a search index over a protein catalogue. Ridge logistic regression and k-means stand in for the study's models and the production clustering tools. Open the live model on its own page ↗
Architecture, stack and core formulation
A protein representation built at scale — cluster, fold, embed and index — and response models that combine species profiles with functional embeddings.
100M+ proteins
Protein sequences from metagenomes, the raw material for functional representation.
Down to ~6M
Sequence and structure clustering collapse redundancy; structure prediction supports function inference.
Protein language models
Each cluster representative embedded with protein language models; about 60× faster than the baseline pipeline.
Vector search
A vector database searching about 6 million sequences in 1–2 seconds; UMAP for exploration.
Response and candidates
Species-level profiles and functional contig embeddings feed random forests, a microbiome taxonomic language model and ensembles.
| Layer | Technology | What it does here |
|---|---|---|
| Clustering | MMseqs2 (sequence), Foldseek (structure) | 100M+ sequences → about 6M clusters |
| Structure | AlphaFold, ESMFold | Structures for function inference and structural clustering |
| Embedding | ESM2, gLM protein language models | Fixed-length functional representations |
| Search | Chroma DB vector store, UMAP | 1–2 s lookups over about 6M sequences |
| Prediction | Random forest, microbiome taxonomic language model, BMCARP, ensembles | Responder vs non-responder across cohorts |
cluster MMseqs2 / Foldseek: sequences (structures) → representatives c, |C| ≈ 6M embed e_c = PLM(seq_c) ∈ ℝ^d ESM2, gLM search q → top-k over the vector index by cosine similarity features sample s → [ species profile_s ; functional embeddings_s ] predict P(responder | s) = RF / language model / BMCARP / ensemble evaluate train on some cohorts, test on others (live model: leave one cohort out)
- Scale first. Clustering before embedding is what made 100M+ sequences tractable — and the clusters double as the search index.
- Function travels better than species. Different species in different populations carry the same functional genes, so functional features generalize across cohorts.
- Candidates, not just scores. Features that stay strong across cohorts are biomarker candidates; the structure was designed for discovery as well as prediction.
| Component | In production | In the live model above |
|---|---|---|
| Representation | ESM2 / gLM embeddings over ~6M clusters | 16-dimensional generated embeddings of 12,000 proteins |
| Clustering | MMseqs2 / Foldseek | k-means with 200 clusters |
| Search | Chroma DB, 1–2 s over ~6M | Cluster-centroid index over 12,000 vectors |
| Models | Random forest, MTLM, BMCARP, ensembles | Ridge logistic regression |
| Data | Public metagenomic cohorts (942 samples, nine cohorts in the SITC study) | 942 generated samples in nine cohorts |
Design notes
Test on cohorts the model has never seen
A biomarker is useful only if it works in the next hospital. Cross-validation that mixes samples from every cohort rewards memorising each cohort's quirks; holding out a whole cohort at a time is the honest test, and the one on which species-level signatures most often fail.
What microbes do travels better than who they are
Different species in different populations can carry the same functional genes. Mapping each sample onto protein-family clusters — functional units found in embedding space — gives features that mean the same thing in every cohort, and candidates that stay strong across folds rather than in one cohort only.
One structure, two jobs
The clustering that makes a hundred million sequences tractable is also what makes them searchable: comparing a query with cluster centres first, then only with the members of the nearest clusters, cuts the work by orders of magnitude at little cost in recall. That is why the embedding, clustering and search work came before the prediction model rather than after it.
Outcome
The work led to a SITC 2024 poster presentation, patent preparation, press materials and follow-up collaboration discussions with external research organizations. It also helped CJ Bioscience and the AI Center register a formal collaborative KPI project.
Limitations
- Association is not mechanism: biomarker candidates from observational cohorts need experimental validation before any clinical use.
- Public cohorts differ in sequencing, processing and response definitions; part of the cross-cohort gap is technical rather than biological.
- The live model's species, proteins, embeddings and responses are generated with a planted functional signal, so it shows the mechanism of the approach, not the size of the real effect.
About the demo and confidentiality
Everything in the embedded model is generated. No patient, cohort, sequence, model or result from CJ Bioscience or the AI Center appears here beyond the figures stated in my CV and the public SITC 2024 poster.