Retrieved Sequence Augmentation for Protein Representation Learning
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Lu, Qi Liu, Sheng Wang, Lingpeng Kong
Abstract
Protein Language Models traditionally depend on Multiple Sequence Alignments (MSA) to incorporate evolutionary knowledge. However, MSA-based approaches suffer from substantial computational overhead and generally underperform in generalizing to de novo proteins. This study reevaluates the role of MSA, proposing it as a retrieval augmentation method and questioning the necessity of sequence alignment. We show that a simple alternative, Retrieved Sequence Augmentation (RSA), can enhance protein representation learning without the need for alignment and cumbersome preprocessing. RSA surpasses MSA Transformer by an average of 5% in both structural and property prediction tasks while being 373 times faster. Additionally, RSA demonstrates enhanced transferability for predicting de novo proteins. This methodology addresses a critical need for efficiency in protein prediction and can be rapidly employed to identify homologous sequences, improve representation learning, and enhance the capacity of Large Language Models to interpret protein structures. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02f95d50-dd64-43ba-9763-d15c2d9d4d22Cited by top-tier papers5
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Rethinking Text-based Protein Understanding: Retrieval or LLM?Juntong Wu, Zijing Liu, He Cao, Li Hao et al.EMNLP 2025 · 7 citations
- Enhancing Protein Mutation Effect Prediction through a Retrieval-Augmented FrameworkRuihan Guo, Rui Wang, Ruidong Wu, Zhizhou Ren et al.NeurIPS 2024 · 3 citations
- Protriever: End-to-End Differentiable Protein Homology Search for Fitness PredictionRuben Weitzman, Peter Mørch Groth, Lood van Niekerk, Aoi Otani et al.ICML 2025
- Retrieval Augmented Zero-Shot Enzyme Generation for Specified SubstrateJiahe Du, Kaixiong Zhou, Xinyu Hong, Zhaozhuo Xu et al.ICML 2025
Builds on13
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
Related papers
- MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure PredictionsLe Zhang, Jiayang Chen, Tao Shen, Yu Li et al.NeurIPS 2024 · 4 citations
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 96 citations
- RMSAGen: Integrating Multiple Sequence Alignment for Function RNA DesignJiyue Jiang, Yanyu Chen, Qingchuan Zhang, Jiayi Li et al.AAAI 2026 · 1 citation
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
