Gut-microbiome signatures of response to immune checkpoint inhibitors tend to fall apart when a model trained on some cohorts meets another: different countries, different species. Describing what the microbes can do — protein families found by embedding and clustering their genes — carries across cohorts better than which species happen to be present. The same clusters, used as an index, make a large protein catalogue searchable in a fraction of the comparisons.
Every gene in the catalogue becomes a protein embedding — ESM2 and gLM in the project, generated vectors here — and the vectors are clustered. In production more than 100 million sequences collapsed into about 6 million clusters with MMseqs2 and Foldseek; here 12,000 proteins form 200 clusters by k-means.
Each sample's species abundances are mapped onto the clusters its species carry, giving a functional profile. A model is trained on eight cohorts and tested on the ninth, for every cohort in turn — the honest test for a biomarker meant to work in a new hospital.
Features that stay among the strongest in every fold are biomarker candidates, not artefacts of one cohort. The cluster centroids double as an index: a query is compared with the centroids first, then only with the members of the nearest few clusters.