A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to Applications
Zhenyu Tao, Wei Xu, Xiaohu You
摘要
The bisimulation metric (BSM) is a powerful tool for computing state similarities within a Markov decision process (MDP), revealing that states closer in BSM have more similar optimal value functions. While BSM has been successfully utilized in reinforcement learning (RL) for tasks like state representation learning and policy exploration, its application to multiple-MDP scenarios, such as policy transfer, remains challenging. Prior work has attempted to generalize BSM to pairs of MDPs, but a lack of rigorous analysis of its mathematical properties has limited further theoretical progress. In this work, we formally establish a generalized bisimulation metric (GBSM) between pairs of MDPs, which is rigorously proven with the three fundamental properties: GBSM symmetry, inter-MDP triangle inequality, and the distance bound on identical state spaces. Leveraging these properties, we theoretically analyse policy transfer, state aggregation, and sampling-based estimation in MDPs, obtaining explicit bounds that are strictly tighter than those derived from the standard BSM. Additionally, GBSM provides a closed-form sample complexity for estimation, improving upon existing asymptotic results based on BSM. Numerical results validate our theoretical findings and demonstrate the effectiveness of GBSM in multi-MDP scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 被引用 171 次
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal 等ICLR 2021 · 被引用 77 次
- Towards Robust Bisimulation Metric LearningMete Kemertas, Tristan Aumentado-ArmstrongNeurIPS 2021 · 被引用 68 次
- Bisimulation Makes Analogies in Goal-Conditioned Reinforcement LearningPhilippe Hansen-Estruch, Amy Zhang, Ashvin Nair, Patrick Yin 等ICML 2022 · 被引用 39 次
- Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation MetricYiming Wang, Ming Yang, Renzhi Dong, Binbin Sun 等NeurIPS 2023 · 被引用 27 次
相关 Paper
- When Is Generalizable Reinforcement Learning Tractable?Dhruv Malik, Yuanzhi Li, Pradeep RavikumarNeurIPS 2021 · 被引用 32 次
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement LearningMenglong Zhang, Fuyuan Qian, Quanying LiuICLR 2025
- Bisimulation Metric for Model Predictive ControlYutaka Shimizu, Masayoshi TomizukaICLR 2025
- Continuous MDP Homomorphisms and Homomorphic Policy GradientSahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger 等NeurIPS 2022 · 被引用 34 次
- Learning Robust State Abstractions for Hidden-Parameter Block MDPsAmy Zhang, Shagun Sodhani, Khimya Khetarpal, Joelle PineauICLR 2021 · 被引用 5 次
