Learning the Minimum Action Distance
Lorenzo Steccanella, Joshua B. Evans, Özgür Şimşek, Anders Jonsson
摘要
This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories, requiring neither reward signals nor the actions executed by the agent. We propose learning the (MAD), defined as the minimum number of actions required to transition between states, as a fundamental metric that captures the underlying structure of an environment. The MAD naturally enables critical downstream tasks such as goal-conditioned reinforcement learning and reward shaping by providing a dense, geometrically meaningful measure of progress. Our self-supervised learning approach constructs an embedding space where the distances between embedded state pairs correspond to their MAD, accommodating both symmetric and asymmetric approximations. We evaluate the framework on a comprehensive suite of environments with known MAD values, encompassing both deterministic and stochastic transition dynamics, discrete and continuous state spaces, and environments with noisy observations. Empirical results show that the proposed approach learns MAD representations more efficiently than existing methods, produces more accurate estimates of the true MAD, and improves performance on downstream goal-reaching tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill DiscoveryKristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, Sergey LevineICLR 2020 · 被引用 94 次
相关 Paper
- MICo: Improved representations via sampling-based state similarity for Markov decision processesPablo Samuel Castro, Tyler Kastner, Prakash Panangaden, Mark RowlandNeurIPS 2021 · 被引用 66 次
- Dynamics-Aware EmbeddingsWilliam F. Whitney, Rajat Agarwal, Kyunghyun Cho, Abhinav GuptaICLR 2020
- Model-Based Visual Planning with Self-Supervised Functional DistancesStephen Tian, Suraj Nair, Frederik Ebert, Sudeep Dasari 等ICLR 2021 · 被引用 69 次
- Reachability-Aware Laplacian Representation in Reinforcement LearningKaixin Wang, Kuangqi Zhou, Jiashi Feng, Bryan Hooi 等ICML 2023 · 被引用 10 次
- Adversarial Intrinsic Motivation for Reinforcement LearningIshan Durugkar, Mauricio Tec, Scott Niekum, Peter StoneNeurIPS 2021 · 被引用 61 次
