Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning
Vittorio Giammarino, Ruiqi Ni, Ahmed H. Qureshi
Abstract
Offline Goal-Conditioned Reinforcement Learning (GCRL) holds great promise for domains such as autonomous navigation and locomotion, where collecting interactive data is costly and unsafe. However, it remains challenging in practice due to the need to learn from datasets with limited coverage of the state-action space and to generalize across long-horizon tasks. To improve on these challenges, we propose a Physics-informed (Pi) regularized loss for value learning, derived from the Eikonal Partial Differential Equation (PDE) and which induces a geometric inductive bias in the learned value function. Unlike generic gradient penalties that are primarily used to stabilize training, our formulation is grounded in continuous-time optimal control and encourages value functions to align with cost-to-go structures. The proposed regularizer is broadly compatible with temporal-difference-based value learning and can be integrated into existing Offline GCRL algorithms. When combined with Hierarchical Implicit Q-Learning (HIQL), the resulting method, Eikonal-regularized HIQL (Eik-HIQL), yields significant improvements in both performance and generalization, with pronounced gains in stitching regimes and large-scale navigation tasks. Code is available at link 1 .
While gradient norm penalties, closely related to the Eikonal PDE residual, have been employed in generative modeling [14] and, more recently, in model-based RL to regularize Q-functions and mitigate overfitting [15], their application in GCRL remains, to our knowledge, unexplored. In contrast to these prior methods, which primarily aim to stabilize training, our regularizer is designed to inject a structural inductive bias into the GCVF, thereby improving sample efficiency and generalization. To the best of our knowledge, this work presents the first use of the Eikonal PDE as a regularization objective in value-based RL, and its first practical deployment in the Offline GCRL setting. More broadly, there has been a growing interest in incorporating physical priors and geometric, distance-like structures into RL algorithms, especially in model-based or Koopmaninspired frameworks [16,17]. These approaches typically leverage the Koopman operator, which assumes access to (or approximations of) the underlying system dynamics, and may incorporate structural information such as reversibility or symmetry of the dynamics. In contrast, our method operates directly on the GCVF in a fully model-free setting.
In parallel, the use of non-Pi constraints for value learning has been extensively studied in Offline RL. Many approaches constrain learned policies to remain close to the behavior policy, either through explicit density modeling [18][19][20] or implicit divergence constraints [21,22]. Others directly regularize the Q-function to assign low values to out-of-distribution actions and improve robustness [23,24].
Our work is also closely related to the literature on GCRL [3,4]. Eik-HIQL extends HIQL [10] by integrating the Eikonal regularizer into the GCVF estimation loss. HIQL itself combines hierarchical actors [25,26] with Implicit Q-Learning [23]. Other GCRL methods include hindsight relabeling [27], contrastive representation learning [13], state-occupancy matching [28], and quasimetric RL [12]. Offline GCRL has also been studied through the lens of goal-conditioned supervised learning (GCSL) [29,30], where goal-reaching policies are trained via conditional imitation or regression. Recent work has analyzed the out-of-distribution goal generalization problem [31] and addressed GCSL via self-supervised reward shaping [32] and occupancy-based score modeling [33].
Beyond RL, Pi losses and neural networks (NNs) have been widely applied to learn parameterizations of PDEs such as Burgers, Schrödinger, and Navier-Stokes equations [34,35]. These methods leverage automatic differentiation to estimate derivatives with respect to NN inputs and solve high-dimensional PDEs. The Eikonal PDE, in particular, has been used in seismology [36] and motion planning [37,38], where distance fields provide essential geometric structure. In this work, we extend the use of the Eikonal PDE to the Offline GCRL domain, demonstrating how its distance-preserving properties enhance GCVF estimation, especially in large environments and when data stitching is required.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e90c043-51dd-445d-a20a-c0a831a98a3cCited by top-tier papers6
- Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement LearningVittorio Giammarino, Ahmed Hussain QureshiICLR 2026 · 7 citations
- Test-Time Graph Search for Goal-Conditioned Reinforcement LearningEvgenii Opryshko, Junwei Quan, Claas Voelcker, Yilun Du et al.ICML 2026 · 6 citations
- Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement LearningJunseok Kim, Dohyeong Kim, Mineui Hong, Songhwai OhICML 2026 · 1 citation
- Action-Sufficient Goal RepresentationsJinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup MoonICML 2026 · 1 citation
- Latent Representation Alignment for Offline Goal-Conditioned Reinforcement LearningHyungkyu Kang, Byeongchan Kim, Min-hwan OhICML 2026
Builds on20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
Related papers
- Conservative Offline Goal-Conditioned Implicit V-LearningKaiqiang Ke, Qian Lin, Zongkai Liu, Shenghong He et al.ICML 2025
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 22 citations
- Continuity-Regularized Flow Matching for Offline Reinforcement LearningXiaocong Chen, Siyu Wang, Lina YaoICML 2026
- Scaling Goal-conditioned Reinforcement Learning with Multistep Quasimetric DistancesBill Zheng, Vivek Myers, Benjamin Eysenbach, Sergey LevineICLR 2026 · 1 citation
- Provably Efficient Offline Goal-Conditioned Reinforcement Learning with General Function Approximation and Single-Policy ConcentrabilityHanlin Zhu, Amy ZhangNeurIPS 2023 · 7 citations
