Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization
Jiajun Fan, Shuaike Shen, Chaoran Cheng, Yuxin Chen, Chumeng Liang, Ge Liu
摘要
Recent advancements in reinforcement learning (RL) have achieved great success in fine-tuning diffusion-based generative models. However, fine-tuning continuous flow-based generative models to align with arbitrary user-defined reward functions remains challenging, particularly due to issues such as policy collapse from overoptimization and the prohibitively high computational cost of likelihoods in continuous-time flows. In this paper, we propose an easy-to-use and theoretically sound RL fine-tuning method, which we term Online Reward-Weighted Conditional Flow Matching with Wasserstein-2 Regularization (ORW-CFM-W2). Our method integrates RL into the flow matching framework to fine-tune generative models with arbitrary reward functions, without relying on gradients of rewards or filtered datasets. By introducing an online reward-weighting mechanism, our approach guides the model to prioritize high-reward regions in the data manifold. To prevent policy collapse and maintain diversity, we incorporate Wasserstein-2 (W2) distance regularization into our method and derive a tractable upper bound for it in flow matching, effectively balancing exploration and exploitation of policy optimization. We provide theoretical analyses to demonstrate the convergence properties and induced data distributions of our method, establishing connections with traditional RL algorithms featuring Kullback-Leibler (KL) regularization and offering a more comprehensive understanding of the underlying mechanisms and learning behavior of our approach. Extensive experiments on tasks including target image generation, image compression, and text-image alignment demonstrate the effectiveness of our method, where our method achieves optimal policy convergence while allowing controllable trade-offs between reward maximization and diversity preservation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningTonghe Zhang, Chao Yu, Sichang Su, Yu WangNeurIPS 2025 · 被引用 101 次
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion ModelsShuchen Xue, Chongjian GE, Shilong Zhang, Yichen Li 等ICML 2026 · 被引用 43 次
- Flow-Based Policy for Online Reinforcement LearningLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun 等NeurIPS 2025 · 被引用 39 次
- Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement LearningChubin Chen, Sujie Hu, Jiashu Zhu, Meiqi Wu 等CVPR 2026 · 被引用 28 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative ModelsJiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen 等NeurIPS 2025 · 被引用 7 次
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu 等NeurIPS 2023 · 被引用 372 次
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-FunctionHyeongyu Kang, Jaewoo Lee, Woocheol Shin, Kiyoung Om 等ICLR 2026 · 被引用 5 次
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-TuningRiccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh, Zebang Shen 等NeurIPS 2025 · 被引用 17 次
- Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion ModelsMin Cheng, Fatemeh Doudi, Dileep Kalathil, Mohammad Ghavamzadeh 等ICLR 2026 · 被引用 6 次
