Metric Residual Network for Sample Efficient Goal-Conditioned Reinforcement Learning
Bo Liu, Yihao Feng, Qiang Liu, Peter Stone
Abstract
Goal-conditioned reinforcement learning (GCRL) has a wide range of potential real-world applications, including manipulation and navigation problems in robotics. Especially in such robotics tasks, sample efficiency is of the utmost importance for GCRL since, by default, the agent is only rewarded when it reaches its goal. While several methods have been proposed to improve the sample efficiency of GCRL, one relatively under-studied approach is the design of neural architectures to support sample efficiency. In this work, we introduce a novel neural architecture for GCRL that achieves significantly better sample efficiency than the commonlyused monolithic network architecture. The key insight is that the optimal action-value function Q * (s, a, g) must satisfy the triangle inequality in a specific sense. Furthermore, we introduce the metric residual network (MRN) that deliberately decomposes the action-value function Q(s, a, g) into the negated summation of a metric plus a residual asymmetric component. MRN provably approximates any optimal actionvalue function Q * (s, a, g), thus making it a fitting neural architecture for GCRL. We conduct comprehensive experiments across 12 standard benchmark environments in GCRL. The empirical results demonstrate that MRN uniformly outperforms other state-of-the-art GCRL neural architectures in terms of sample efficiency. The code is publicly available at https://github.com/Cranial-XIX/metric-residual-network .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-MakingVivek Myers, Chongyi Zheng, Anca D. Dragan, Sergey Levine et al.ICML 2024 · 38 citations
- Offline Goal-conditioned Reinforcement Learning with Quasimetric RepresentationsVivek Myers, Bill Zheng, Benjamin Eysenbach, Sergey LevineNeurIPS 2025 · 26 citations
- Inference via Interpolation: Contrastive Representations Provably Enable Planning and InferenceBenjamin Eysenbach, Vivek Myers, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2024 · 23 citations
- Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction FollowingVivek Myers, Bill Zheng, Anca D. Dragan, Kuan Fang et al.NeurIPS 2025 · 13 citations
- Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement LearningVittorio Giammarino, Ahmed Hussain QureshiICLR 2026 · 7 citations
Related papers
- Optimal Goal-Reaching Reinforcement Learning via Quasimetric LearningTongzhou Wang, Antonio Torralba, Phillip Isola, Amy ZhangICML 2023 · 88 citations
- An Inductive Bias for Distances: Neural Nets that Respect the Triangle InequalitySilviu Pitis, Harris Chan, Kiarash Jamali, Jimmy BaICLR 2020 · 31 citations
- Scalable Multi-Agent Reinforcement Learning through Intelligent Information AggregationSiddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan et al.ICML 2023 · 73 citations
- -Equivariant Reinforcement LearningDian Wang, Robin Walters, Robert PlattICLR 2022 · 105 citations
- Scaling Goal-conditioned Reinforcement Learning with Multistep Quasimetric DistancesBill Zheng, Vivek Myers, Benjamin Eysenbach, Sergey LevineICLR 2026 · 1 citation
