Diffusion on Language Model Encodings for Protein Sequence Generation
Viacheslav Meshchaninov, Pavel V. Strashnov, Andrey Shevtsov, Fedor Nikolaev, Nikita Ivanisenko, Olga L. Kardymon, Dmitry P. Vetrov
摘要
Protein design necessitates a profound understanding of the intricate nature of the protein universe. While approaches based on discrete diffusion and autoregression are actively developing in the field of protein sequence generation, continuous diffusion remains underappreciated and underexplored. To address this gap, this research introduces DiMA, a latent diffusion model that leverages Gaussian diffusion on representations derived from protein language models,such as ESM-2 and CHEAP, to generate amino acid sequences. We quantitatively investigate the impact of various components of the latent diffusion model and protein encoders, revealing their contributions to enhanced protein generation performance. Additionally, we conduct an extensive evaluation of existing methods alongside DiMA using multiple metrics across two protein modalities, covering quality, novelty, diversity, and distribution matching of generated proteins. Our findings demonstrate that DiMA consistently produces novel, high-quality, and diverse protein sequences that accurately reflect the inherent structural and functional diversity of the protein space. Furthermore, we show that the proposed model can be easily adapted to address conditional tasks, such as protein family generation and inpainting. This work advances the field of protein design by providing a robust framework for latent diffusion on various protein representations, facilitating high-quality protein sequence generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo 等NeurIPS 2025 · 被引用 13 次
- GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein GenerationPavel V. Strashnov, Andrey Shevtsov, Viacheslav Meshchaninov, Olga L. Kardymon 等ICLR 2026
- MMCP-GEN: A Modality-Extensible Diffusion Language Model for Conditional Protein Sequence GenerationZeyu An, Wanyu Lin, Feng Tan, Shujun WangCVPR 2026
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
相关 Paper
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- ProteinAE: Protein Diffusion Autoencoders for Structure EncodingShaoning Li, Le Zhuo, Yusong Wang, Mingyu Li 等ICLR 2026 · 被引用 3 次
- Latent Diffusion for Language GenerationJustin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman 等NeurIPS 2023 · 被引用 177 次
- Bridging Protein Sequences and Microscopy Images with Unified Diffusion ModelsDihan Zheng, Bo HuangICML 2025
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICLR 2025
