Efficient Exploration in Resource-Restricted Reinforcement Learning
Zhihai Wang, Taoxing Pan, Qi Zhou, Jie Wang
Abstract
In many real-world applications of reinforcement learning (RL), performing actions requires consuming certain types of resources that are non-replenishable in each episode. Typical applications include robotic control with limited energy and video games with consumable items. In tasks with non-replenishable resources, we observe that popular RL methods such as soft actor critic suffer from poor sample efficiency. The major reason is that, they tend to exhaust resources fast and thus the subsequent exploration is severely restricted due to the absence of resources. To address this challenge, we first formalize the aforementioned problem as a resource-restricted reinforcement learning, and then propose a novel resource-aware exploration bonus (RAEB) to make reasonable usage of resources. An appealing feature of RAEB is that, it can significantly reduce unnecessary resource-consuming trials while effectively encouraging the agent to explore unvisited states. Experiments demonstrate that the proposed RAEB significantly outperforms state-of-the-art exploration strategies in resource-restricted reinforcement learning environments, improving the sample efficiency by up to an order of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- ChiPFormer: Transferable Chip Placement via Offline Decision TransformerYao Lai, Jinxin Liu, Zhentao Tang, Bin Wang et al.ICML 2023 · 69 citations
- Reinforcement Learning within Tree Search for Fast Macro PlacementZijie Geng, Jie Wang, Ziyan Liu, Siyuan Xu et al.ICML 2024 · 23 citations
- State Sequences Prediction via Fourier Transform for Representation LearningMingxuan Ye, Yufei Kuang, Jie Wang, Rui Yang et al.NeurIPS 2023 · 18 citations
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo et al.ICML 2024 · 17 citations
- Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation FrameworkZhihai Wang, Jie Wang, Qingyue Yang, Yinqi Bai et al.NeurIPS 2024 · 17 citations
Builds on2
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu et al.NeurIPS 2021 · 106 citations
Related papers
- KEA: Keeping Exploration Alive by Proactively Coordinating Exploration StrategiesShih-Min Yang, Martin Magnusson, Johannes A. Stork, Todor StoyanovICML 2025
- Modeling Human Exploration Through Resource-Rational Reinforcement LearningMarcel Binz, Eric SchulzNeurIPS 2022 · 23 citations
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin et al.ICML 2024 · 8 citations
- Sample-Efficient Multiagent Reinforcement Learning with Reset ReplayYaodong Yang, Guangyong Chen, Jianye Hao, Pheng-Ann HengICML 2024 · 9 citations
- SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic LearningJinyang Wu, Shuo Yang, Yuhao Shen, Shuai Zhang et al.ACL 2026 · 9 citations
