Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
Jiaqi Han, Austin Wang, Minkai Xu, Wenda Chu, Meihua Dang, Haotian Ye, Huayu Chen, Yisong Yue, Stefano Ermon
摘要
Discrete diffusion models have demonstrated great promise in modeling various sequence data, ranging from human language to biological sequences. Inspired by the success of RL in language models, there is growing interest in further improving the models by alignment with a certain reward. In this work, we propose an offline preference optimization method to approach trajectory alignment for discrete diffusion models. Instead of applying the reward on the final output and backpropagating the gradient to the entire denoising process, we decompose the problem into a set of stepwise alignment objectives by matching the per-step posterior. This framework enables efficient diffusion optimization, is compatible with arbitrary reward functions, and importantly, yields an equivalent optimal solution under additive factorization of the trajectory reward. Experiments across multiple domains including DNA sequence design, protein inverse folding, and language modeling consistently demonstrate the superiority of our approach. Notably, it achieves an up to 12% improvement over the most competitive RL-based baseline in terms of predicted activity on DNA sequence design, and further improves the GSM8K score from 78.6 to 81.2 on LLaDA-8B-Instruct for language modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language ModelsXiaohang Tang, Rares Dolga, Sangwoong Yoon, Ilija BogunovicICLR 2026 · 被引用 70 次
- Inpainting-Guided Policy Optimization for Diffusion Large Language ModelsSiyan Zhao, Mengchen Liu, Jing Huang, Miao Liu 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein DesignChenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang 等ICLR 2025
- Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior PredictionJarrid Rector-Brooks, Mohsin Hasan, Zhangzhi Peng, Cheng-Hao Liu 等ICLR 2025
- Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Ehsan Hajiramezanali, Gabriele Scalia 等NeurIPS 2024 · 被引用 31 次
- Aligning Diffusion Behaviors with Q-functions for Efficient Continuous ControlHuayu Chen, Kaiwen Zheng, Hang Su, Jun ZhuNeurIPS 2024 · 被引用 13 次
- Structure-based RNA Design by Step-wise Optimization of Latent Diffusion ModelQi Si, Xuyang Liu, Penglei Wang, Xin Guo 等AAAI 2026
