FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling
Jianwei Zhu, Yu Shi, Ran Bi, Peiran Jin, Chang Liu, Zhe Zhang, Haitao Huang, Zekun Guo, Pipi Hu, Fusong Ju, Lin Huang, Xinwei Tai
摘要
Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-based language models (e.g., ESM-2) capture sequence semantics at scale, and a number of recent works incorporate structural signals into sequence encoders. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting evolutionary couplings, but their reliance on homologous sequences makes them less reliable in highly mutated or alignment-sparse regimes. We present FlexRibbon ‡ , a pretrained protein model that jointly learns from amino acid sequences and three-dimensional structures. Our pretraining strategy combines masked language modeling with diffusionbased denoising, enabling bidirectional sequence-structure learning without requiring MSAs. Trained on both experimentally resolved structures and AlphaFold 2 predictions, FlexRibbon captures global folds as well as flexible conformations critical for biological function. Evaluated across diverse tasks spanning interface design, intermolecular interaction prediction, and protein function prediction, FlexRibbon establishes new state-of-the-art performance on 12 different tasks, with particularly strong gains in mutation-rich settings where MSA-based methods often struggle.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- GeoDiff: A Geometric Diffusion Model for Molecular Conformation GenerationMinkai Xu, Lantao Yu, Yang Song, Chence Shi 等ICLR 2022 · 被引用 695 次
- Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models for Protein StructuresShitong Luo, Yufeng Su, Xingang Peng, Sheng Wang 等NeurIPS 2022 · 被引用 331 次
相关 Paper
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICLR 2025
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl 等ICLR 2025
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- Rigidity-Aware Geometric Pretraining for Protein Design and Conformational EnsemblesZhanghan Ni, Yanjing Li, Zeju Qiu, Bernhard Schölkopf 等ICLR 2026 · 被引用 2 次
- MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure PredictionsLe Zhang, Jiayang Chen, Tao Shen, Yu Li 等NeurIPS 2024 · 被引用 4 次
