Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design
Chenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang, Avantika Lal, Tommi S. Jaakkola, Sergey Levine, Aviv Regev, Hanchen Wang, Tommaso Biancalani
摘要
Recent studies have demonstrated the strong empirical performance of diffusion models on discrete sequences (i.e., discrete diffusion models) across domains from natural language to biological sequence generation. For example, in the protein inverse folding task, where the goal is to generate a protein sequence from a given backbone structure, conditional diffusion models have achieved impressive results in generating natural-like sequences that fold back into the original structure. However, practical design tasks often require not only modeling a conditional distribution but also optimizing specific task objectives. For instance, in the inverse folding task, we may prefer protein sequences with high stability. To address this, we consider the scenario where we have pre-trained discrete diffusion models that can generate natural-like sequences, as well as reward models that map sequences to task objectives. We then formulate the reward maximization problem within discrete diffusion models, analogous to reinforcement learning (RL), while minimizing the KL divergence against pretrained diffusion models to preserve naturalness. To solve this RL problem, we propose a novel algorithm, DRAKES, that enables direct backpropagation of rewards through entire trajectories generated by diffusion models, by making the originally nondifferentiable trajectories differentiable using the Gumbel-Softmax trick. Our theoretical analysis indicates that our approach can generate sequences that are both natural-like (i.e., have a high probability under a pretrained model) and yield high rewards. While similar tasks have been recently explored in diffusion models for continuous domains, our work addresses unique algorithmic and theoretical challenges specific to discrete diffusion models, which arise from their foundation in continuous-time Markov chains rather than Brownian motion. Finally, we demonstrate the effectiveness of our algorithm in generating DNA and protein sequences that optimize enhancer activity and protein stability, respectively, important tasks for gene therapies and protein-based therapeutics. The code is available at https://github.com/ChenyuWang-Monica/DRAKES .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language ModelsXiaohang Tang, Rares Dolga, Sangwoong Yoon, Ilija BogunovicICLR 2026 · 被引用 70 次
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language ModelsChenyu Wang, Paria Rashidinejad, DiJia Andy Su, Song Jiang 等ICLR 2026 · 被引用 45 次
- Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsOussama Zekri, Nicolas BoulléNeurIPS 2025 · 被引用 42 次
- Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow ModelsYingqing Guo, Yukang Yang, Hui Yuan, Mengdi WangNeurIPS 2025 · 被引用 29 次
- MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal ControlYuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu 等NeurIPS 2025 · 被引用 24 次
它引用的顶会 Paper31
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- Discrete Diffusion Trajectory Alignment via Stepwise DecompositionJiaqi Han, Austin Wang, Minkai Xu, Wenda Chu 等ICLR 2026 · 被引用 11 次
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based DecodingXiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia 等NeurIPS 2025 · 被引用 147 次
- Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior PredictionJarrid Rector-Brooks, Mohsin Hasan, Zhangzhi Peng, Cheng-Hao Liu 等ICLR 2025
- Structure-based RNA Design by Step-wise Optimization of Latent Diffusion ModelQi Si, Xuyang Liu, Penglei Wang, Xin Guo 等AAAI 2026
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 被引用 12 次
