Bio AI · Protein embeddings · Biomarker discoveryCJ AI Center, with CJ BioscienceWrite-up October 2026 · 7 min read

Microbiome signals of immunotherapy response, and a search engine for proteins

Immune checkpoint inhibitors work for some patients and not others, and the gut microbiome is one of the most promising places to look for the difference. This is the work with CJ Bioscience researchers that turned more than 100 million protein sequences into a searchable embedding space and used it to predict response and surface biomarker candidates — presented at SITC 2024.

Built withESM2gLMAlphaFold · ESMFoldMMseqs2FoldseekChroma DBUMAPRandom forestMicrobiome language model
MKTLLVAGG…MSEQRLIAF…MAKVLYDGT…MTRPELSQA…MGKIVLASN…MQHLTAEPR…… 100M+ SEQUENCES embed ≈ 6M CLUSTERS Search: 6M clusters in 1–2 svector store · centroid index Response model: +15% over baselinefunctional features · biomarker candidates
Sequences become embeddings, embeddings become clusters, and the clusters serve twice: as an index for fast search and as functional features for predicting response.

Why the microbiome, and why it is hard

Immune checkpoint inhibitors do not work equally for all patients, and response rates are limited. Predicting response before treatment, and identifying the biomarkers related to it, can support clinical decision-making and drug development. For CJ Bioscience's biomarker platform strategy, that meant AI-based analysis to predict immunotherapy response from microbiome data and to explore biomarker candidates.

Two things make this difficult. Microbiome signatures found in one cohort often fail in another, because populations carry different species. And the underlying data — protein sequences, metagenomes, microbiome profiles — is enormous, so any representation of what the microbes can do must be built and searched at scale.

What was built

Live model, computed in your browser on generated metagenomes the size of the public data used in the SITC poster: 942 samples in nine cohorts, each with its own species mix, and a response that depends on a function carried by different species in different cohorts. Top: a species-only model and a model with protein-family clusters, each trained on eight cohorts and tested on the ninth. Bottom: the same clusters used as a search index over a protein catalogue. Ridge logistic regression and k-means stand in for the study's models and the production clustering tools. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

A protein representation built at scale — cluster, fold, embed and index — and response models that combine species profiles with functional embeddings.

1 · Sequences

100M+ proteins

Protein sequences from metagenomes, the raw material for functional representation.

metagenomes
2 · Cluster

Down to ~6M

Sequence and structure clustering collapse redundancy; structure prediction supports function inference.

MMseqs2FoldseekAlphaFold / ESMFold
3 · Embed

Protein language models

Each cluster representative embedded with protein language models; about 60× faster than the baseline pipeline.

ESM2gLM
4 · Index

Vector search

A vector database searching about 6 million sequences in 1–2 seconds; UMAP for exploration.

Chroma DBUMAP
5 · Predict

Response and candidates

Species-level profiles and functional contig embeddings feed random forests, a microbiome taxonomic language model and ensembles.

random forestMTLMensemble
Stack
LayerTechnologyWhat it does here
ClusteringMMseqs2 (sequence), Foldseek (structure)100M+ sequences → about 6M clusters
StructureAlphaFold, ESMFoldStructures for function inference and structural clustering
EmbeddingESM2, gLM protein language modelsFixed-length functional representations
SearchChroma DB vector store, UMAP1–2 s lookups over about 6M sequences
PredictionRandom forest, microbiome taxonomic language model, BMCARP, ensemblesResponder vs non-responder across cohorts
Core formulation
cluster     MMseqs2 / Foldseek:  sequences (structures) → representatives c,   |C| ≈ 6M
embed       e_c = PLM(seq_c) ∈ ℝ^d                          ESM2, gLM
search      q → top-k over the vector index by cosine similarity
features    sample s → [ species profile_s ; functional embeddings_s ]
predict     P(responder | s) = RF / language model / BMCARP / ensemble
evaluate    train on some cohorts, test on others                      (live model: leave one cohort out)
  • Scale first. Clustering before embedding is what made 100M+ sequences tractable — and the clusters double as the search index.
  • Function travels better than species. Different species in different populations carry the same functional genes, so functional features generalize across cohorts.
  • Candidates, not just scores. Features that stay strong across cohorts are biomarker candidates; the structure was designed for discovery as well as prediction.
In production vs in the live model
ComponentIn productionIn the live model above
RepresentationESM2 / gLM embeddings over ~6M clusters16-dimensional generated embeddings of 12,000 proteins
ClusteringMMseqs2 / Foldseekk-means with 200 clusters
SearchChroma DB, 1–2 s over ~6MCluster-centroid index over 12,000 vectors
ModelsRandom forest, MTLM, BMCARP, ensemblesRidge logistic regression
DataPublic metagenomic cohorts (942 samples, nine cohorts in the SITC study)942 generated samples in nine cohorts

Design notes

Test on cohorts the model has never seen

A biomarker is useful only if it works in the next hospital. Cross-validation that mixes samples from every cohort rewards memorising each cohort's quirks; holding out a whole cohort at a time is the honest test, and the one on which species-level signatures most often fail.

What microbes do travels better than who they are

Different species in different populations can carry the same functional genes. Mapping each sample onto protein-family clusters — functional units found in embedding space — gives features that mean the same thing in every cohort, and candidates that stay strong across folds rather than in one cohort only.

One structure, two jobs

The clustering that makes a hundred million sequences tractable is also what makes them searchable: comparing a query with cluster centres first, then only with the members of the nearest clusters, cuts the work by orders of magnitude at little cost in recall. That is why the embedding, clustering and search work came before the prediction model rather than after it.

Outcome

+15%prediction accuracy over the prior machine-learning baseline
100M → 6Mprotein sequences into clusters; embedding about 60× faster
1–2 ssearch over about 6 million sequences

The work led to a SITC 2024 poster presentation, patent preparation, press materials and follow-up collaboration discussions with external research organizations. It also helped CJ Bioscience and the AI Center register a formal collaborative KPI project.

Limitations

About the demo and confidentiality

Everything in the embedded model is generated. No patient, cohort, sequence, model or result from CJ Bioscience or the AI Center appears here beyond the figures stated in my CV and the public SITC 2024 poster.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterMethod research, large-scale embedding and clustering, vector search, response modelling; co-author of the SITC 2024 poster. Demo re-implemented on generated data for this site.