Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
Vivek Myers, Chongyi Zheng, Anca D. Dragan, Sergey Levine, Benjamin Eysenbach
摘要
Temporal distances lie at the heart of many algorithms for planning, control, and reinforcement learning that involve reaching goals, allowing one to estimate the transit time between two states. However, prior attempts to define such temporal distances in stochastic settings have been stymied by an important limitation: these prior approaches do not satisfy the triangle inequality. This is not merely a definitional concern, but translates to an inability to generalize and find shortest paths. In this paper, we build on prior work in contrastive learning and quasimetrics to show how successor features learned by contrastive learning (after a change of variables) form a temporal distance that does satisfy the triangle inequality, even in stochastic settings. Importantly, this temporal distance is computationally efficient to estimate, even in high-dimensional and stochastic settings. Experiments in controlled settings and benchmark suites demonstrate that an RL algorithm based on these new temporal distances exhibits combinatorial generalization (i.e., "stitching") and can sometimes learn more quickly than prior methods, including those based on quasimetrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching CapabilitiesKevin Wang, Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski 等NeurIPS 2025 · 被引用 46 次
- Dual Goal RepresentationsSeohong Park, Deepinder Mann, Sergey LevineICLR 2026 · 被引用 15 次
- Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction FollowingVivek Myers, Bill Zheng, Anca D. Dragan, Kuan Fang 等NeurIPS 2025 · 被引用 13 次
- Transitive RL: Value Learning via Divide and ConquerSeohong Park, Aditya Oberai, Pranav Atreya, Sergey LevineICLR 2026 · 被引用 12 次
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 被引用 9 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Offline Goal-conditioned Reinforcement Learning with Quasimetric RepresentationsVivek Myers, Bill Zheng, Benjamin Eysenbach, Sergey LevineNeurIPS 2025 · 被引用 26 次
- Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic RewardsFaisal Mohamed, Catherine Ji, Benjamin Eysenbach, Glen BersethICLR 2026 · 被引用 1 次
- Scaling Goal-conditioned Reinforcement Learning with Multistep Quasimetric DistancesBill Zheng, Vivek Myers, Benjamin Eysenbach, Sergey LevineICLR 2026 · 被引用 1 次
- Contrastive Difference Predictive CodingChongyi Zheng, Ruslan Salakhutdinov, Benjamin EysenbachICLR 2024 · 被引用 32 次
- Contrastive Representations for Temporal ReasoningAlicja Ziarko, Michal Bortkiewicz, Michal Zawalski, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 8 次
