CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model
Milong Ren, Tian Zhu, Haicang Zhang
摘要
De novo protein design aims to create novel protein structures and sequences unseen in nature. Recent structure-oriented design methods typically employ a two-stage strategy, where structure design and sequence design modules are trained separately, and the backbone structures and sequences are generated sequentially in inference. While diffusion-based generative models like RFdiffusion show great promise in structure design, they face inherent limitations within the two-stage framework. First, the sequence design module risks overfitting as the accuracy of the generated structures may not align with that of the crystal structures used for training. Second, the sequence design module lacks interaction with the structure design module to further optimize the generated structures. To address these challenges, we propose CarbonNovo, a unified energy-based model for jointly generating protein structure and sequence. Specifically, we leverage a score-based generative model and Markov Random Fields for describing the energy landscape of protein structure and sequence. In CarbonNovo, the structure and sequence design module communicates at each diffusion step, encouraging the generation of more coherent structure-sequence pairs. Moreover, the unified framework allows for incorporating the protein language models as evolutionary constraints for generated proteins. The rigorous evaluation demonstrates that CarbonNovo outperforms two-stage methods across various metrics, including designability, novelty, sequence plausibility, and Rosetta energy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach 等NeurIPS 2022 · 被引用 384 次
相关 Paper
- A Joint Diffusion Model with Pre-Trained Priors for RNA Sequence-Structure Co-DesignXiner Li, Masatoshi Uehara, Xingyu Su, Gabriele Scalia 等ICLR 2026
- Protein Sequence and Structure Co-Design with Equivariant TranslationChence Shi, Chuanrui Wang, Jiarui Lu, Bozitao Zhong 等ICLR 2023 · 被引用 13 次
- Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue CloudsYeqing Lin, Mohammed AlQuraishiICML 2023 · 被引用 105 次
- RNAFlow: RNA Structure & Sequence Design via Inverse Folding-Based Flow MatchingDivya Nori, Wengong JinICML 2024 · 被引用 23 次
- Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein DesignYue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen 等ICML 2021 · 被引用 56 次
