MMCP-GEN: A Modality-Extensible Diffusion Language Model for Conditional Protein Sequence Generation
Zeyu An, Wanyu Lin, Feng Tan, Shujun Wang
Abstract
Recent advances in diffusion-based language models (DLMs) have shown remarkable potential for de novo protein design. However, enabling controllable protein generation requires integrating diverse biological conditions, such as structure, functions, and chemical interactions, each represented in distinct modalities. Existing approaches often either support a single condition or treat multiple conditions through separate modality-specific encoders. This isolation limits cross-modal interaction, reduces generation quality, and complicates the incorporation of new conditions without retraining or redesigning the backbone. To address these limitations, we introduce MMCP-GEN, a DLM for Multi-Modal, Multi-Condition Protein sequence GENeration. MMCP-GEN establishes a new paradigm for controllable protein generation under complex multimodal constraints. Its core is a modality-composable and extensible conditioning mechanism that fuses heterogeneous biological conditions via learnable queries and modality-indicator heads, enabling disentangled, extensible, and cross-modal condition integration without retraining the backbone. A joint generation-and-scoring objective further aligns sequence recovery with structural fidelity. Empirically, MMCP-GEN achieves state-of-the-art performance across structure-, function-, and ligand-conditioned tasks, improving sequence recovery by up to 5% and outperforming attentive baselines in diverse functional annotation tasks. These results establish MMCP-GEN as a general and high-fidelity framework for controllable protein generation. The source code is publicly available at https://github.com/WanyuGroup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8cc30fa-d8d3-4db9-8b27-721a3134f108Builds on11
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- SE(3) diffusion model with application to protein backbone generationJason Yim, Brian L. Trippe, Valentin De Bortoli, Emile Mathieu et al.ICML 2023 · 313 citations
- Protein Design with Guided Discrete DiffusionNate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner et al.NeurIPS 2023 · 246 citations
- SE(3)-Stochastic Flow Matching for Protein Backbone GenerationAvishek Joey Bose, Tara Akhound-Sadegh, Guillaume Huguet, Kilian Fatras et al.ICLR 2024 · 162 citations
- Structure-informed Language Models Are Protein DesignersZaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou et al.ICML 2023 · 130 citations
Related papers
- CFP-Gen: Combinatorial Functional Protein Generation via Diffusion Language ModelsJunbo Yin, Chao Zha, Wenjia He, Chencheng Xu et al.ICML 2025
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICLR 2025
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- A Multi-Modal Contrastive Diffusion Model for Therapeutic Peptide GenerationYongkang Wang, Xuan Liu, Feng Huang, Zhankun Xiong et al.AAAI 2024 · 28 citations
- Co-Generative De Novo Functional Protein DesignXinRui Chen, YIZHEN LUO, Siqi Fan, Zaiqing NieICML 2026
