Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation Metric
Yiming Wang, Ming Yang, Renzhi Dong, Binbin Sun, Furui Liu, Leong Hou U
摘要
Reward shaping is an effective technique for integrating domain knowledge into reinforcement learning (RL). However, traditional approaches like potential-based reward shaping totally rely on manually designing shaping reward functions, which significantly restricts exploration efficiency and introduces human cognitive biases. While a number of RL methods have been proposed to boost exploration by designing an intrinsic reward signal as exploration bonus. Nevertheless, these methods heavily rely on the count-based episodic term in their exploration bonus which falls short in scalability. To address these limitations, we propose a general end-to-end potential-based exploration bonus for deep RL via potentials of state discrepancy, which motivates the agent to discover novel states and provides them with denser rewards without manual intervention. Specifically, we measure the novelty of adjacent states by calculating their distance using the bisimulation metric-based potential function, which enhances agent exploration and ensures policy invariance. In addition, we offer a theoretical guarantee on our inverse dynamic bisimulation metric, bounding the value difference and ensuring that the agent explores states with higher TD error, thus significantly improving training efficiency. The proposed approach is named LIBERTY (expLoration vIa Bisimulation mEtRic-based sTate discrepancY) which is comprehensively evaluated on the MuJoCo and the Arcade Learning Environments. Extensive experiments have verified the superiority and scalability of our algorithm compared with other competitive methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan 等ICLR 2026 · 被引用 32 次
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin 等ICLR 2026 · 被引用 28 次
- Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration BonusYiming Wang, Kaiyan Zhao, Furui Liu, Leong Hou UNeurIPS 2024 · 被引用 15 次
- Bootstrapped Reward ShapingJacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. KulkarniAAAI 2025 · 被引用 7 次
- A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to ApplicationsZhenyu Tao, Wei Xu, Xiaohu YouNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper8
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 被引用 171 次
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu 等NeurIPS 2021 · 被引用 106 次
相关 Paper
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu 等ICML 2026 · 被引用 1 次
- Reward Uncertainty for Exploration in Preference-based Reinforcement LearningXinran Liang, Katherine Shu, Kimin Lee, Pieter AbbeelICLR 2022
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement LearningMenglong Zhang, Fuyuan Qian, Quanying LiuICLR 2025
- MICo: Improved representations via sampling-based state similarity for Markov decision processesPablo Samuel Castro, Tyler Kastner, Prakash Panangaden, Mark RowlandNeurIPS 2021 · 被引用 66 次
- Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICML 2023 · 被引用 17 次
