CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model
Milong Ren, Tian Zhu, Haicang Zhang
Abstract
De novo protein design aims to create novel protein structures and sequences unseen in nature. Recent structure-oriented design methods typically employ a two-stage strategy, where structure design and sequence design modules are trained separately, and the backbone structures and sequences are generated sequentially in inference. While diffusion-based generative models like RFdiffusion show great promise in structure design, they face inherent limitations within the two-stage framework. First, the sequence design module risks overfitting as the accuracy of the generated structures may not align with that of the crystal structures used for training. Second, the sequence design module lacks interaction with the structure design module to further optimize the generated structures. To address these challenges, we propose CarbonNovo, a unified energy-based model for jointly generating protein structure and sequence. Specifically, we leverage a score-based generative model and Markov Random Fields for describing the energy landscape of protein structure and sequence. In CarbonNovo, the structure and sequence design module communicates at each diffusion step, encouraging the generation of more coherent structure-sequence pairs. Moreover, the unified framework allows for incorporating the protein language models as evolutionary constraints for generated proteins. The rigorous evaluation demonstrates that CarbonNovo outperforms two-stage methods across various metrics, including designability, novelty, sequence plausibility, and Rosetta energy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a667f19b-99ef-4d70-87d6-a6a6ed02e0f3Builds on15
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach et al.NeurIPS 2022 · 384 citations
Related papers
- A Joint Diffusion Model with Pre-Trained Priors for RNA Sequence-Structure Co-DesignXiner Li, Masatoshi Uehara, Xingyu Su, Gabriele Scalia et al.ICLR 2026
- Protein Sequence and Structure Co-Design with Equivariant TranslationChence Shi, Chuanrui Wang, Jiarui Lu, Bozitao Zhong et al.ICLR 2023 · 13 citations
- Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue CloudsYeqing Lin, Mohammed AlQuraishiICML 2023 · 105 citations
- RNAFlow: RNA Structure & Sequence Design via Inverse Folding-Based Flow MatchingDivya Nori, Wengong JinICML 2024 · 23 citations
- Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein DesignYue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen et al.ICML 2021 · 56 citations
