floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
Bhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral Kumar
Abstract
A hallmark of modern large-scale machine learning techniques is the use of training objectives that provide dense supervision to intermediate computations, such as teacher forcing the next token in language models or denoising step-by-step in diffusion models. This enables models to learn complex functions in a generalizable manner. Motivated by this observation, we investigate the benefits of iterative computation for temporal difference (TD) methods in reinforcement learning (RL). Typically, they represent value functions in a monolithic fashion, without iterative compute. We introduce floq (flow-matching Q-functions), an approach that parameterizes the Q-function using a velocity field and trains it with techniques from flow-matching, typically used in generative modeling. This velocity field underneath the flow is trained using a TD-learning objective, which bootstraps from values produced by a target velocity field, computed by running multiple steps of numerical integration. Crucially, floq allows for more fine-grained control and scaling of the Q-function capacity than monolithic architectures, by appropriately setting the number of integration steps. Across a suite of challenging offline RL benchmarks and online fine-tuning tasks, floq improves performance by nearly 1.8x. floq scales capacity far better than standard TD-learning architectures, highlighting the potential of iterative computation for value learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf70ae41-ed91-49e4-8a66-cf9b5774bd66Cited by top-tier papers5
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 9 citations
- ReFORM: Reflected Flows for On-support Offline RL via Noise ManipulationSongyuan Zhang, Oswin So, H. M. Sabbir Ahmad, Eric Yang Yu et al.ICLR 2026 · 5 citations
- Reinforcement Learning via Value Gradient FlowHaoran Xu, Kaiwen Hu, Somayeh Sojoudi, Amy ZhangICLR 2026 · 4 citations
- GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow PoliciesHe Zhang, Ying Sun, Hui XiongICLR 2026 · 3 citations
- What Does Flow-Matching Bring to TD-Learning?Bhavya Agrawalla, Michal Nauman, Aviral KumarICML 2026 · 1 citation
Builds on46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan et al.ICML 2026
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement LearningXuan Thanh Nguyen, Chang Dong YooICLR 2026 · 11 citations
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-LearningThanh Nguyen, Tri Ton, Hongbin Choe, Minh-Tung Luu et al.ICML 2026 · 3 citations
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 36 citations
