Reinforcement Learning via Value Gradient Flow
Haoran Xu, Kaiwen Hu, Somayeh Sojoudi, Amy Zhang
Abstract
We study behavior-regularized reinforcement learning (RL), where regularization toward a reference distribution (the dataset in offline RL or the base model in LLM RL finetuning) is essential to prevent value over-optimization caused by erroneous out-of-distribution extrapolation. Existing methods either rely on reparameterized policy gradient, which are difficult to scale to large generative models, or on reject sampling, which can be overly conservative when attempting to move beyond the behavior support. In this paper, we propose Value Gradient Flow (VGF), a scalable new paradigm for behavior-regularized RL. VGF casts behavior-regularized RL as an optimal transport problem that maps the reference distribution to the valueinduced optimal policy distribution. We solve this transport problem via discrete gradient flow, where value gradients guide particles initialized from the reference distribution. Our analysis shows that VGF imposes regularization implicitly by controlling the transport budget. VGF eliminates explicit policy parameterization while remaining expressive and flexible, this enables adaptive test-time scaling by adjusting the transport budget. Extensive experiments demonstrate that VGF significantly outperforms prior methods, achieving state-of-the-art results on offline RL benchmarks (D4RL, OGBench) and LLM RL tasks. Code and runs can be found at https://ryanxhr.github.io/vgf ' '() * ๐ฃ(๐ ! " , ๐ก) โ โ # ! " log ๐ * (๐ ! " |๐ ) ๐ ! " ๐ !%& " ๐ * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38ea0098-bde0-4030-a391-1266ed9cb38cBuilds on33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 ยท 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 ยท 24,707 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 ยท 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 ยท 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 ยท 1,402 citations
Related papers
- Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement LearningFranki Nguimatsia Tiofack, Thรฉotime Le Hellard, Fabian Schramm, Nicolas Perrin-Gilbert et al.ICLR 2026 ยท 8 citations
- Flow Actor-Critic for Offline Reinforcement LearningJongseong Chae, Jongeui Park, Yongjae Shin, Gyeongmin Kim et al.ICLR 2026 ยท 7 citations
- Behavior Regularization with Flow Latent Policy for Offline Reinforcement LearningYulong Xia, Fuchun SunAAAI 2026
- Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement LearningChen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao et al.ICML 2025
- REG: In-Sample RL via Regularizing the Evaluation GapHanpu Shen, Weining Shen, Roy FoxICML 2026
