Scalable Methods for Computing State Similarity in Deterministic Markov Decision Processes
Pablo Samuel Castro
Abstract
We present new algorithms for computing and approximating bisimulation metrics in Markov Decision Processes (MDPs). Bisimulation metrics are an elegant formalism that capture behavioral equivalence between states and provide strong theoretical guarantees on differences in optimal behaviour. Unfortunately, their computation is expensive and requires a tabular representation of the states, which has thus far rendered them impractical for large problems. In this paper we present a new version of the metric that is tied to a behavior policy in an MDP, along with an analysis of its theoretical properties. We then present two new algorithms for approximating bisimulation metrics in large, deterministic MDPs. The first does so via sampling and is guaranteed to converge to the true metric. The second is a differentiable loss which allows us to learn an approximation even for continuous state MDPs, which prior to this work had not been possible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b16b9634-0d2a-4989-bb69-b6d853d36a94Cited by top-tier papers79
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
- Reinforcement Learning with Action-Free Pre-Training from VideosYounggyo Seo, Kimin Lee, Stephen James, Pieter AbbeelICML 2022 · 150 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 125 citations
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement LearningTianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu et al.NeurIPS 2020 · 112 citations
Related papers
- A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to ApplicationsZhenyu Tao, Wei Xu, Xiaohu YouNeurIPS 2025 · 6 citations
- Compositional Behavioral Semantics for State Abstraction in Reinforcement LearningYivan Zhang, Ziyan Luo, Manuel BaltieriICML 2026
- BeigeMaps: Behavioral Eigenmaps for Reinforcement Learning from ImagesSandesh Adhikary, Anqi Li, Byron BootsICML 2024 · 1 citation
- Approximate Probabilistic Bisimulation for Continuous-Time Markov ChainsTimm Spork, Christel Baier, Joost-Pieter Katoen, Sascha Klüppelholz et al.CAV 2025 · 1 citation
- Metrics and Continuity in Reinforcement LearningCharline Le Lan, Marc G. Bellemare, Pablo Samuel CastroAAAI 2021 · 40 citations
