One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
Zeyuan Wang, Da Li, Yulin Chen, Ye Shi, Liang Bai, Tianyuan Yu, Yanwei Fu
摘要
We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with Q-learning. While one-step Gaussian policies enable fast inference, they struggle to capture complex, multimodal action distributions. Existing flow-based methods improve expressivity but typically rely on distillation and two-stage training when trained with Q-learning. To overcome these limitations, we propose to reformulate MeanFlow to enable direct noise-to-action generation by integrating the velocity field and noise-to-action transformation into a single policy network-eliminating the need for separate velocity estimation. We explore several reformulation variants and identify an effective residual formulation that supports expressive and stable policy learning. Our method offers three key advantages: 1) efficient one-step noise-to-action generation, 2) expressive modelling of multimodal action distributions, and 3) efficient and stable policy learning via Q-learning in a single-stage training setup. Extensive experiments on 73 tasks across the OG-Bench and D4RL benchmarks demonstrate that our method achieves strong performance in both offline and offline-toonline reinforcement learning settings. Code is available at https://github.com/HiccupRL/MeanFlowQL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-LearningThanh Nguyen, Tri Ton, Hongbin Choe, Minh-Tung Luu 等ICML 2026 · 被引用 3 次
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan 等ICML 2026
- Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-LearningSungyoung Lee, Dohyeong Kim, Eshan Balachandar, Zelal Mustafaoglu 等ICML 2026
它引用的顶会 Paper31
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Mean Flows for One-step Generative ModelingZhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter 等NeurIPS 2025 · 被引用 628 次
相关 Paper
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement LearningXuan Thanh Nguyen, Chang Dong YooICLR 2026 · 被引用 11 次
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Flow Actor-Critic for Offline Reinforcement LearningJongseong Chae, Jongeui Park, Yongjae Shin, Gyeongmin Kim 等ICLR 2026 · 被引用 7 次
- Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action GenerationGuojian Zhan, Letian Tao, Pengcheng Wang, Yixiao Wang 等ICLR 2026 · 被引用 11 次
- Flow-Based Single-Step Completion for Efficient and Expressive Policy LearningPrajwal Koirala, Cody FlemingICLR 2026 · 被引用 12 次
