Embedding-perturbed Exploration Preference Optimization for Flow Models
Sujie Hu, Chubin Chen, Jiashu Zhu, Jiahong Wu, Xiangxiang Chu, Xiu Li
Abstract
Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization frameworks (e.g., GRPO) face a critical limitation: the rapid decay of intra-group variance. As the distinctiveness among samples within a group diminishes, the variance approaches zero. This eliminates the very learning signal required for optimization, rendering the process unstable and forcing the policy into premature stagnation or reward hacking. Existing strategies, such as varying the initial noise or increasing group sizes, often fail to address this fundamental issue, resulting in training instability or diminishing returns. To overcome these challenges, we propose mbedding-perturbed xploration Preference Optimization (PO), a novel framework that sustains optimization through embedding-level perturbation. Our method introduces structured, embedding-level perturbations within sample groups, guaranteeing a robust variance that preserves the discriminative signal throughout the training process. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art baselines, achieving a more faithful alignment with human preference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2ee5df9-a224-40f3-8cd6-243d2f0583a8Cited by top-tier papers1
Ask how each one uses itBuilds on56
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
Related papers
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive GenerationShikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu et al.CVPR 2026 · 2 citations
- Seeing What Matters: Visual Preference Policy Optimization for Visual GenerationZiqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou et al.CVPR 2026 · 9 citations
- Preference-Enhanced Reinforcement Learning for Pluralistic Image InpaintingPeng Zhou, Muqi Huang, Tianshuo Qu, Jingyang Wang et al.ICML 2026
- Group Robust Preference Optimization in Reward-free RLHFShyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta et al.NeurIPS 2024 · 122 citations
- Fine-Grained GRPO for Precise Preference Alignment in Flow ModelsYujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang et al.CVPR 2026 · 19 citations
