floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
Bhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral Kumar
摘要
A hallmark of modern large-scale machine learning techniques is the use of training objectives that provide dense supervision to intermediate computations, such as teacher forcing the next token in language models or denoising step-by-step in diffusion models. This enables models to learn complex functions in a generalizable manner. Motivated by this observation, we investigate the benefits of iterative computation for temporal difference (TD) methods in reinforcement learning (RL). Typically, they represent value functions in a monolithic fashion, without iterative compute. We introduce floq (flow-matching Q-functions), an approach that parameterizes the Q-function using a velocity field and trains it with techniques from flow-matching, typically used in generative modeling. This velocity field underneath the flow is trained using a TD-learning objective, which bootstraps from values produced by a target velocity field, computed by running multiple steps of numerical integration. Crucially, floq allows for more fine-grained control and scaling of the Q-function capacity than monolithic architectures, by appropriately setting the number of integration steps. Across a suite of challenging offline RL benchmarks and online fine-tuning tasks, floq improves performance by nearly 1.8x. floq scales capacity far better than standard TD-learning architectures, highlighting the potential of iterative computation for value learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 被引用 9 次
- ReFORM: Reflected Flows for On-support Offline RL via Noise ManipulationSongyuan Zhang, Oswin So, H. M. Sabbir Ahmad, Eric Yang Yu 等ICLR 2026 · 被引用 5 次
- Reinforcement Learning via Value Gradient FlowHaoran Xu, Kaiwen Hu, Somayeh Sojoudi, Amy ZhangICLR 2026 · 被引用 4 次
- GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow PoliciesHe Zhang, Ying Sun, Hui XiongICLR 2026 · 被引用 3 次
- What Does Flow-Matching Bring to TD-Learning?Bhavya Agrawalla, Michal Nauman, Aviral KumarICML 2026 · 被引用 1 次
它引用的顶会 Paper46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan 等ICML 2026
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement LearningXuan Thanh Nguyen, Chang Dong YooICLR 2026 · 被引用 11 次
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-LearningThanh Nguyen, Tri Ton, Hongbin Choe, Minh-Tung Luu 等ICML 2026 · 被引用 3 次
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
