MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training
Bo Chen, Zhilei Bei, Xingyi Cheng, Pan Li, Jie Tang, Le Song
Abstract
Multiple Sequence Alignment (MSA) plays a pivotal role in unveiling the evolutionary trajectories of protein families. The accuracy of protein structure predictions is often compromised for protein sequences that lack sufficient homologous information to construct high-quality MSA. Although various methods have been proposed to generate virtual MSA under these conditions, they fall short in comprehensively capturing the intricate co-evolutionary patterns within MSA or require guidance from external oracle models. Here we introduce MSAGPT, a novel approach to prompt protein structure predictions via MSA generative pre-training in the low-MSA regime. MSAGPT employs a simple yet effective 2D evolutionary positional encoding scheme to model the complex evolutionary patterns. Endowed by this, its flexible 1D MSA decoding framework facilitates zero-or few-shot learning. More-over, we demonstrate that leveraging the feedback from AlphaFold2 can further enhance the model’s capacity via Rejective Fine-tuning (RFT) and Reinforcement Learning from AF2 Feedback (RLAF). Extensive experiments confirm the efficacy of MSAGPT in generating faithful virtual MSA to enhance the structure prediction accuracy (up to +8.5% TM-Score on few-shot scenarios). The transfer learning capabilities also highlight its great potential for facilitating other protein tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2cb239d-f854-413a-9dd3-ad02b754e79bCited by top-tier papers4
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin et al.EMNLP 2024 · 4 citations
- ProteinAE: Protein Diffusion Autoencoders for Structure EncodingShaoning Li, Le Zhuo, Yusong Wang, Mingyu Li et al.ICLR 2026 · 3 citations
- Rigidity-Aware Geometric Pretraining for Protein Design and Conformational EnsemblesZhanghan Ni, Yanjing Li, Zeju Qiu, Bernhard Schölkopf et al.ICLR 2026 · 2 citations
- Steering Protein Family Design through Profile Bayesian FlowJingjing Gong, Yu Pei, Siyu Long, Yuxuan Song et al.ICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier et al.ICML 2021 · 686 citations
Related papers
- MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure PredictionsLe Zhang, Jiayang Chen, Tao Shen, Yu Li et al.NeurIPS 2024 · 4 citations
- FlexRibbon: Joint Sequence and Structure Pretraining for Protein ModelingJianwei Zhu, Yu Shi, Ran Bi, Peiran Jin et al.ICLR 2026 · 2 citations
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen et al.ICML 2024 · 10 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan et al.AAAI 2026 · 1 citation
