Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning
Vittorio Giammarino, Ruiqi Ni, Ahmed H. Qureshi
摘要
Offline Goal-Conditioned Reinforcement Learning (GCRL) holds great promise for domains such as autonomous navigation and locomotion, where collecting interactive data is costly and unsafe. However, it remains challenging in practice due to the need to learn from datasets with limited coverage of the state-action space and to generalize across long-horizon tasks. To improve on these challenges, we propose a Physics-informed (Pi) regularized loss for value learning, derived from the Eikonal Partial Differential Equation (PDE) and which induces a geometric inductive bias in the learned value function. Unlike generic gradient penalties that are primarily used to stabilize training, our formulation is grounded in continuous-time optimal control and encourages value functions to align with cost-to-go structures. The proposed regularizer is broadly compatible with temporal-difference-based value learning and can be integrated into existing Offline GCRL algorithms. When combined with Hierarchical Implicit Q-Learning (HIQL), the resulting method, Eikonal-regularized HIQL (Eik-HIQL), yields significant improvements in both performance and generalization, with pronounced gains in stitching regimes and large-scale navigation tasks. Code is available at link 1 .
While gradient norm penalties, closely related to the Eikonal PDE residual, have been employed in generative modeling [14] and, more recently, in model-based RL to regularize Q-functions and mitigate overfitting [15], their application in GCRL remains, to our knowledge, unexplored. In contrast to these prior methods, which primarily aim to stabilize training, our regularizer is designed to inject a structural inductive bias into the GCVF, thereby improving sample efficiency and generalization. To the best of our knowledge, this work presents the first use of the Eikonal PDE as a regularization objective in value-based RL, and its first practical deployment in the Offline GCRL setting. More broadly, there has been a growing interest in incorporating physical priors and geometric, distance-like structures into RL algorithms, especially in model-based or Koopmaninspired frameworks [16,17]. These approaches typically leverage the Koopman operator, which assumes access to (or approximations of) the underlying system dynamics, and may incorporate structural information such as reversibility or symmetry of the dynamics. In contrast, our method operates directly on the GCVF in a fully model-free setting.
In parallel, the use of non-Pi constraints for value learning has been extensively studied in Offline RL. Many approaches constrain learned policies to remain close to the behavior policy, either through explicit density modeling [18][19][20] or implicit divergence constraints [21,22]. Others directly regularize the Q-function to assign low values to out-of-distribution actions and improve robustness [23,24].
Our work is also closely related to the literature on GCRL [3,4]. Eik-HIQL extends HIQL [10] by integrating the Eikonal regularizer into the GCVF estimation loss. HIQL itself combines hierarchical actors [25,26] with Implicit Q-Learning [23]. Other GCRL methods include hindsight relabeling [27], contrastive representation learning [13], state-occupancy matching [28], and quasimetric RL [12]. Offline GCRL has also been studied through the lens of goal-conditioned supervised learning (GCSL) [29,30], where goal-reaching policies are trained via conditional imitation or regression. Recent work has analyzed the out-of-distribution goal generalization problem [31] and addressed GCSL via self-supervised reward shaping [32] and occupancy-based score modeling [33].
Beyond RL, Pi losses and neural networks (NNs) have been widely applied to learn parameterizations of PDEs such as Burgers, Schrödinger, and Navier-Stokes equations [34,35]. These methods leverage automatic differentiation to estimate derivatives with respect to NN inputs and solve high-dimensional PDEs. The Eikonal PDE, in particular, has been used in seismology [36] and motion planning [37,38], where distance fields provide essential geometric structure. In this work, we extend the use of the Eikonal PDE to the Offline GCRL domain, demonstrating how its distance-preserving properties enhance GCVF estimation, especially in large environments and when data stitching is required.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement LearningVittorio Giammarino, Ahmed Hussain QureshiICLR 2026 · 被引用 7 次
- Test-Time Graph Search for Goal-Conditioned Reinforcement LearningEvgenii Opryshko, Junwei Quan, Claas Voelcker, Yilun Du 等ICML 2026 · 被引用 6 次
- Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement LearningJunseok Kim, Dohyeong Kim, Mineui Hong, Songhwai OhICML 2026 · 被引用 1 次
- Action-Sufficient Goal RepresentationsJinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup MoonICML 2026 · 被引用 1 次
- Latent Representation Alignment for Offline Goal-Conditioned Reinforcement LearningHyungkyu Kang, Byeongchan Kim, Min-hwan OhICML 2026
它引用的顶会 Paper20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
相关 Paper
- Conservative Offline Goal-Conditioned Implicit V-LearningKaiqiang Ke, Qian Lin, Zongkai Liu, Shenghong He 等ICML 2025
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 被引用 22 次
- Continuity-Regularized Flow Matching for Offline Reinforcement LearningXiaocong Chen, Siyu Wang, Lina YaoICML 2026
- Scaling Goal-conditioned Reinforcement Learning with Multistep Quasimetric DistancesBill Zheng, Vivek Myers, Benjamin Eysenbach, Sergey LevineICLR 2026 · 被引用 1 次
- Provably Efficient Offline Goal-Conditioned Reinforcement Learning with General Function Approximation and Single-Policy ConcentrabilityHanlin Zhu, Amy ZhangNeurIPS 2023 · 被引用 7 次
