TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
Yihong Luo, Tianyang Hu, Weijian Luo, Jing Tang
Abstract
While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models remain an unsolved problem. Existing RL approaches for few-step diffusion models strongly rely on back-propagating through differentiable reward models, thereby excluding the majority of important real-world reward signals, e.g., non-differentiable rewards such as humans' binary likeness, object counts, etc. To properly incorporate non-differentiable rewards to improve few-step generative models, we introduce TDM-R1, a novel reinforcement learning paradigm built upon a leading few-step model, Trajectory Distribution Matching (TDM). TDM-R1 decouples the learning process into surrogate reward learning and generator learning. Furthermore, we developed practical methods to obtain per-step reward signals along the deterministic generation trajectory of TDM, resulting in a unified RL post-training method that significantly improves few-step models' ability with generic rewards. We conduct extensive experiments ranging from textrendering, visual quality, and preference alignment. All results demonstrate that TDM-R1 is a powerful reinforcement learning paradigm for few-step text-to-image models, achieving state-of-the-art reinforcement learning performances on both in-domain and out-of-domain metrics. Furthermore, TDM-R1 also scales effectively to the recent strong Z-Image model, consistently outperforming both its 100-NFE and few-step variants with only 4 NFEs. Project page: https://github.com/Luo-Yihong/TDM-R1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cc4aba0-1d71-4041-8dee-bbf9e9bbc9cbBuilds on49
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image GenerationYihong Luo, Tianyang Hu, Weijian Luo, Kenji Kawaguchi et al.NeurIPS 2025 · 20 citations
- Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion ModelsXiaole Xian, Xilin He, Wenting Chen, Wenshuang Liu et al.ICLR 2026
- Curriculum Direct Preference Optimization for Diffusion and Consistency ModelsFlorinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe et al.CVPR 2025
- A Dense Reward View on Aligning Text-to-Image Diffusion with PreferenceShentao Yang, Tianqi Chen, Mingyuan ZhouICML 2024 · 53 citations
- Learning Few-Step Diffusion Models by Trajectory Distribution MatchingYihong Luo, Tianyang Hu, Jiacheng Sun, Yujun Cai et al.ICCV 2025 · 3 citations
