Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
Ruben Weitzman, Peter Mørch Groth, Lood van Niekerk, Aoi Otani, Yarin Gal, Debora Susan Marks, Pascal Notin
Abstract
Retrieving homologous protein sequences is essential for a broad range of protein modeling tasks such as fitness prediction, protein design, structure modeling, and protein-protein interactions. Traditional workflows have relied on a two-step process: first retrieving homologs via Multiple Sequence Alignments (MSA), then training models on one or more of these alignments. However, MSA-based retrieval is computationally expensive, struggles with highly divergent sequences or complex insertions & deletions patterns, and operates independently of the downstream modeling objective. We introduce Protriever, an end-to-end differentiable framework that learns to retrieve relevant homologs while simultaneously training for the target task. When applied to protein fitness prediction, Protriever achieves state-of-the-art performance compared to sequence-based models that rely on MSA-based homolog retrieval, while being two orders of magnitude faster through efficient vector search. Protriever is both architectureand task-agnostic, and can flexibly adapt to different retrieval strategies and protein databases at inference time -offering a scalable alternative to alignment-centric approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3317248c-987e-43ef-8fbb-3b8bc06a5fd3Builds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier et al.ICML 2021 · 686 citations
Related papers
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin et al.EMNLP 2024 · 4 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- Data Distillation for extrapolative protein design through exact preference optimizationMostafa Karimi, Sharmi Banerjee, Tommi S. Jaakkola, Bella Dubrov et al.ICLR 2025
- FlexRibbon: Joint Sequence and Structure Pretraining for Protein ModelingJianwei Zhu, Yu Shi, Ran Bi, Peiran Jin et al.ICLR 2026 · 2 citations
- Fast End-to-End Learning on Protein SurfacesFreyr Sverrisson, Jean Feydy, Bruno E. Correia, Michael M. BronsteinCVPR 2021
