Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion Models
Xiaole Xian, Xilin He, Wenting Chen, Wenshuang Liu, Wenqi Mu, Yancheng He, Liang Li, Yi Zhang, Xiangyu Yue
Abstract
Recent advances in diffusion models for visual generation have sparked interest in human preference alignment, similar to developments in Large Language Models. While reward model (RM) based approaches enable trajectory-aware optimization by evaluating intermediate timesteps, they face two critical challenges: unreliable reward estimation on noisy latents due to pixel-level models' sensitivity to noise interference, and single-timestep preference evaluation across sampling trajectories where single-timestep evaluations can yield inconsistent preference rankings depending on the selected timestep. To address these limitations, we propose a comprehensive framework with targeted solutions for each challenge. To achieve noise compatibility for reliable reward estimation, we introduce the Score-based Latent Reward Model (SLRM), which leverages the complete diffusion model as a preference discriminator with learnable task tokens and a score enhancement mechanism that explicitly preserves noise compatibility by augmenting preference logits with the denoising score function. To ensure consistent preference evaluation across trajectories, we develop Trajectory Advantages Preference Optimization (TAPO), which strategically performs Stochastic Differential Equations sampling and reward evaluation at multiple timesteps to dynamically capture trajectory advantages while identifying preference inconsistencies and preventing erroneous trajectory selection. Extensive experiments on Text-to-Image and Textto-Video generation tasks demonstrate significant improvements on noisy latent evaluation and alignment performance. The code is available at TAPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationTao Zhang, Cheng Da, Kun Ding, Huan Yang et al.NeurIPS 2025 · 38 citations
- DSPO: Direct Score Preference Optimization for Diffusion Model AlignmentHuaisheng Zhu, Teng Xiao, Vasant G. HonavarICLR 2025
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 4 citations
- Rethinking Direct Preference Optimization in Diffusion ModelsJunyong Kang, Seohyun Lim, Kyungjune Baek, Hyunjung ShimAAAI 2026
- Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human PreferencesYunhong Lu, Qichao Wang, Hengyuan Cao, Xiaoyin Xu et al.ICML 2025
