Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning
Mingqi Yuan, Bo Li, Xin Jin, Wenjun Zeng
Abstract
We present AIRS: Automatic Intrinsic Reward Shaping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task return in real-time, providing reliable exploration incentives and alleviating the biased objective problem. Moreover, we develop an intrinsic reward toolkit 1 to provide efficient and reliable implementations of diverse intrinsic reward approaches. We test AIRS on various tasks of MiniGrid, Procgen, and DeepMind Control Suite. Extensive simulation demonstrates that AIRS can outperform the benchmarking schemes and achieve superior performance with simple architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9639224-b5ea-4a05-94dc-4b8faa3b76ddCited by top-tier papers8
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu et al.NeurIPS 2024 · 14 citations
- Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement LearningXinran Li, Zifan Liu, Shibo Chen, Jun ZhangICML 2024 · 11 citations
- Exploratory Diffusion Model for Unsupervised Reinforcement LearningChengyang Ying, Huayu Chen, Xinning Zhou, Zhongkai Hao et al.ICLR 2026 · 4 citations
- Action-Dependent Optimality-Preserving Reward ShapingGrant C. Forbes, Jianxun Wang, Leonardo Villalobos-Arias, Arnav Jhala et al.ICML 2025
- Generative Modeling of Discrete Latent Structures via Dynamic Policy GradientsStefan Ivanovic, Ge Liu, Mohammed El-KebirICML 2026
Builds on9
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
Related papers
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 48 citations
- Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse RewardsRati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2022 · 122 citations
- Revisiting Intrinsic Reward for Exploration in Procedurally Generated EnvironmentsKaixin Wang, Kuangqi Zhou, Bingyi Kang, Jiashi Feng et al.ICLR 2023
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade et al.NeurIPS 2022 · 115 citations
- Exploit Reward Shifting in Value-Based Deep-RL: Optimistic Curiosity-Based Exploration and Conservative Exploitation via Linear Reward ShapingHao Sun, Lei Han, Rui Yang, Xiaoteng Ma et al.NeurIPS 2022 · 46 citations
