Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation Metric
Yiming Wang, Ming Yang, Renzhi Dong, Binbin Sun, Furui Liu, Leong Hou U
Abstract
Reward shaping is an effective technique for integrating domain knowledge into reinforcement learning (RL). However, traditional approaches like potential-based reward shaping totally rely on manually designing shaping reward functions, which significantly restricts exploration efficiency and introduces human cognitive biases. While a number of RL methods have been proposed to boost exploration by designing an intrinsic reward signal as exploration bonus. Nevertheless, these methods heavily rely on the count-based episodic term in their exploration bonus which falls short in scalability. To address these limitations, we propose a general end-to-end potential-based exploration bonus for deep RL via potentials of state discrepancy, which motivates the agent to discover novel states and provides them with denser rewards without manual intervention. Specifically, we measure the novelty of adjacent states by calculating their distance using the bisimulation metric-based potential function, which enhances agent exploration and ensures policy invariance. In addition, we offer a theoretical guarantee on our inverse dynamic bisimulation metric, bounding the value difference and ensuring that the agent explores states with higher TD error, thus significantly improving training efficiency. The proposed approach is named LIBERTY (expLoration vIa Bisimulation mEtRic-based sTate discrepancY) which is comprehensively evaluated on the MuJoCo and the Arcade Learning Environments. Extensive experiments have verified the superiority and scalability of our algorithm compared with other competitive methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1213ff1e-98a4-4238-b8e8-abfa6e9828cdCited by top-tier papers14
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan et al.ICLR 2026 · 32 citations
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin et al.ICLR 2026 · 28 citations
- Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration BonusYiming Wang, Kaiyan Zhao, Furui Liu, Leong Hou UNeurIPS 2024 · 15 citations
- Bootstrapped Reward ShapingJacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. KulkarniAAAI 2025 · 7 citations
- A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to ApplicationsZhenyu Tao, Wei Xu, Xiaohu YouNeurIPS 2025 · 6 citations
Builds on8
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu et al.NeurIPS 2021 · 106 citations
Related papers
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu et al.ICML 2026 · 1 citation
- Reward Uncertainty for Exploration in Preference-based Reinforcement LearningXinran Liang, Katherine Shu, Kimin Lee, Pieter AbbeelICLR 2022
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement LearningMenglong Zhang, Fuyuan Qian, Quanying LiuICLR 2025
- MICo: Improved representations via sampling-based state similarity for Markov decision processesPablo Samuel Castro, Tyler Kastner, Prakash Panangaden, Mark RowlandNeurIPS 2021 · 66 citations
- Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICML 2023 · 17 citations
