Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
Shuchen Xue, Chongjian GE, Shilong Zhang, Yichen Li, Zhi-Ming Ma
摘要
Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where pre-training and RL post-training share the same log-likelihood formulation. In contrast, recent RL approaches for diffusion models, most notably Denoising Diffusion Policy Optimization (DDPO), optimize an objective different from the pretraining objectives-score/flow matching loss. In this work, we establish a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence. Building on this analysis, we introduce Advantage Weighted Matching (AWM), a policy-gradient method for diffusion. It uses the same score/flow-matching loss as pretraining to obtain a lower-variance objective and reweights each sample by its advantage. In effect, AWM raises the influence of high-reward samples and suppresses low-reward ones while keeping the modeling objective identical to pretraining. This unifies pretraining and RL conceptually and practically, is consistent with policy-gradient theory, reduces variance, and yields faster convergence. This simple yet effective design yields substantial benefits: on GenEval, OCR, and PickScore benchmarks, AWM delivers up to a 24× speedup over Flow-GRPO (which builds on DDPO), when applied to Stable Diffusion 3.5 Medium and FLUX, without compromising generation quality. Code is available at https://github.com/scxue/advantage_weighted_matching .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss DesignJaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk 等ICML 2026 · 被引用 15 次
- LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step TrajectoriesZhanhao Liang, Tao Yang, Jie Wu, Chengjian Feng 等CVPR 2026 · 被引用 6 次
- TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable RewardYihong Luo, Tianyang Hu, Weijian Luo, Jing TangICML 2026 · 被引用 5 次
- Diffusion Alignment as Variational Expectation-MaximizationJaewoo Lee, Minsu Kim, Sanghyeok Choi, Inhyuck Song 等ICLR 2026 · 被引用 2 次
- FAIL: Flow Matching Adversarial Imitation Learning for Image GenerationYeyao Ma, Chen Li, Xiaosong Zhang, Han Hu 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Reinforcing Diffusion Models by Direct Group Preference OptimizationYihong Luo, Tianyang Hu, Jing TangICLR 2026 · 被引用 13 次
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelKai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge 等CVPR 2024 · 被引用 34 次
- DSPO: Direct Score Preference Optimization for Diffusion Model AlignmentHuaisheng Zhu, Teng Xiao, Vasant G. HonavarICLR 2025
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim 等ICLR 2026 · 被引用 103 次
