MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training
Bo Chen, Zhilei Bei, Xingyi Cheng, Pan Li, Jie Tang, Le Song
摘要
Multiple Sequence Alignment (MSA) plays a pivotal role in unveiling the evolutionary trajectories of protein families. The accuracy of protein structure predictions is often compromised for protein sequences that lack sufficient homologous information to construct high-quality MSA. Although various methods have been proposed to generate virtual MSA under these conditions, they fall short in comprehensively capturing the intricate co-evolutionary patterns within MSA or require guidance from external oracle models. Here we introduce MSAGPT, a novel approach to prompt protein structure predictions via MSA generative pre-training in the low-MSA regime. MSAGPT employs a simple yet effective 2D evolutionary positional encoding scheme to model the complex evolutionary patterns. Endowed by this, its flexible 1D MSA decoding framework facilitates zero-or few-shot learning. More-over, we demonstrate that leveraging the feedback from AlphaFold2 can further enhance the model’s capacity via Rejective Fine-tuning (RFT) and Reinforcement Learning from AF2 Feedback (RLAF). Extensive experiments confirm the efficacy of MSAGPT in generating faithful virtual MSA to enhance the structure prediction accuracy (up to +8.5% TM-Score on few-shot scenarios). The transfer learning capabilities also highlight its great potential for facilitating other protein tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin 等EMNLP 2024 · 被引用 4 次
- ProteinAE: Protein Diffusion Autoencoders for Structure EncodingShaoning Li, Le Zhuo, Yusong Wang, Mingyu Li 等ICLR 2026 · 被引用 3 次
- Rigidity-Aware Geometric Pretraining for Protein Design and Conformational EnsemblesZhanghan Ni, Yanjing Li, Zeju Qiu, Bernhard Schölkopf 等ICLR 2026 · 被引用 2 次
- Steering Protein Family Design through Profile Bayesian FlowJingjing Gong, Yu Pei, Siyu Long, Yuxuan Song 等ICLR 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier 等ICML 2021 · 被引用 686 次
相关 Paper
- MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure PredictionsLe Zhang, Jiayang Chen, Tao Shen, Yu Li 等NeurIPS 2024 · 被引用 4 次
- FlexRibbon: Joint Sequence and Structure Pretraining for Protein ModelingJianwei Zhu, Yu Shi, Ran Bi, Peiran Jin 等ICLR 2026 · 被引用 2 次
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen 等ICML 2024 · 被引用 10 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan 等AAAI 2026 · 被引用 1 次
