Value Flows
Perry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh, Benjamin Eysenbach
Abstract
While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods exploit the return distribution to provide stronger learning signals and to enable applications in exploration and safe RL. While the predominant method for estimating the return distribution is by modeling it as a categorical distribution over discrete bins or estimating a finite number of quantiles, such approaches leave unanswered questions about the fine-grained structure of the return distribution and about how to distinguish states with high return uncertainty for decision-making. The key idea in this paper is to use modern, flexible flow-based models to estimate the full future return distributions and identify those states with high return variance. We do so by formulating a new flow-matching objective that generates probability density paths satisfying the distributional Bellman equation. Building upon the learned flow models, we estimate the return uncertainty of distinct states using a new flow derivative ODE. We additionally use this uncertainty information to prioritize learning a more accurate return estimation on certain transitions. We compare our method (Value Flows) with prior methods in the offline and online-to-online settings. Experiments on state-based and image-based benchmark tasks demonstrate that Value Flows achieves a improvement on average in success rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh et al.ICML 2026 · 7 citations
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 1 citation
- SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online TransferNathan S. de Lara, Florian ShkurtiICML 2026
- Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-LearningSungyoung Lee, Dohyeong Kim, Eshan Balachandar, Zelal Mustafaoglu et al.ICML 2026
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelJing Zhang, Linjiajie Fang, Kexin Shi, Wenjia Wang et al.NeurIPS 2024 · 14 citations
- Path-Coupled Bellman Flows for Distributional Reinforcement LearningBoyang Xu, Qing Zou, Siqin Yang, Hao YanICML 2026
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Value Diffusion Reinforcement LearningXiaoliang Hu, Fuyun Wang, Tong Zhang, Zhen CuiNeurIPS 2025 · 2 citations
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan et al.ICML 2026
