MICo: Improved representations via sampling-based state similarity for Markov decision processes
Pablo Samuel Castro, Tyler Kastner, Prakash Panangaden, Mark Rowland
Abstract
We present a new behavioural distance over the state space of a Markov decision process, and demonstrate the use of this distance as an effective means of shaping the learnt representations of deep reinforcement learning agents. While existing notions of state similarity are typically difficult to learn at scale due to high computational cost and lack of sample-based algorithms, our newly-proposed distance addresses both of these issues. In addition to providing detailed theoretical analysis, we provide empirical evidence that learning this distance alongside the value function yields structured and informative representations, including strong results on the Arcade Learning Environment benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abd14fc0-495a-48fc-bbc4-08986cc90b7cCited by top-tier papers12
- SimSR: Simple Distance-Based State Representations for Deep Reinforcement LearningHongyu Zang, Xin Li, Mingzhong WangAAAI 2022 · 20 citations
- Understanding and Addressing the Pitfalls of Bisimulation-based Representations in Offline Reinforcement LearningHongyu Zang, Xin Li, Leiji Zhang, Yang Liu et al.NeurIPS 2023 · 15 citations
- The Curse of Diversity in Ensemble-Based ExplorationZhixuan Lin, Pierluca D'Oro, Evgenii Nikishin, Aaron C. CourvilleICLR 2024 · 9 citations
- Focus-Then-Decide: Segmentation-Assisted Reinforcement LearningChao Chen, Jiacheng Xu, Weijian Liao, Hao Ding et al.AAAI 2024 · 7 citations
- Policy-Independent Behavioral Metric-Based Representation for Deep Reinforcement LearningWeijian Liao, Zongzhang Zhang, Yang YuAAAI 2023 · 7 citations
Builds on10
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill et al.ICML 2020 · 153 citations
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 125 citations
Related papers
- BeigeMaps: Behavioral Eigenmaps for Reinforcement Learning from ImagesSandesh Adhikary, Anqi Li, Byron BootsICML 2024 · 1 citation
- Learning Representations via a Robust Behavioral Metric for Deep Reinforcement LearningJianda Chen, Sinno Jialin PanNeurIPS 2022 · 19 citations
- Learning the Minimum Action DistanceLorenzo Steccanella, Joshua B. Evans, Özgür Şimşek, Anders JonssonICML 2026
- Learning Generalizable Representations for Reinforcement Learning via Adaptive Meta-learner of Behavioral SimilaritiesJianda Chen, Sinno Jialin PanICLR 2022 · 6 citations
- Reachability-Aware Laplacian Representation in Reinforcement LearningKaixin Wang, Kuangqi Zhou, Jiashi Feng, Bryan Hooi et al.ICML 2023 · 10 citations
