Task-agnostic Exploration in Reinforcement Learning
Xuezhou Zhang, Yuzhe Ma, Adish Singla
Abstract
Efficient exploration is one of the main challenges in reinforcement learning (RL). Most existing sample-efficient algorithms assume the existence of a single reward function during exploration. In many practical scenarios, however, there is not a single underlying reward function to guide the exploration, for instance, when an agent needs to learn many skills simultaneously, or multiple conflicting objectives need to be balanced. To address these challenges, we propose the task-agnostic RL framework: In the exploration phase, the agent first collects trajectories by exploring the MDP without the guidance of a reward function. After exploration, it aims at finding near-optimal policies for tasks, given the collected trajectories augmented with sampled rewards for each task. We present an efficient task-agnostic RL algorithm, UCBZero, that finds -optimal policies for arbitrary tasks after at most exploration episodes. We also provide an lower bound, showing that the dependency on is unavoidable. Furthermore, we provide an -independent sample complexity bound of UCBZero in the statistically easier setting when the ground truth reward functions are known.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 686df2ee-2024-41f7-bec1-ba7ebb1db479Cited by top-tier papers27
- A Sharp Analysis of Model-based Reinforcement Learning with Self-PlayQinghua Liu, Tiancheng Yu, Yu Bai, Chi JinICML 2021 · 137 citations
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann et al.ICML 2021 · 110 citations
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 62 citations
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 60 citations
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 48 citations
Builds on1
Related papers
- Near Optimal Reward-Free Reinforcement LearningZihan Zhang, Simon S. Du, Xiangyang JiICML 2021 · 15 citations
- Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse TasksZiping Xu, Zifan Xu, Runxuan Jiang, Peter Stone et al.ICLR 2024 · 2 citations
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 5 citations
- Interesting Object, Curious Agent: Learning Task-Agnostic ExplorationSimone Parisi, Victoria Dean, Deepak Pathak, Abhinav GuptaNeurIPS 2021 · 58 citations
- On Reward-Free Reinforcement Learning with Linear Function ApproximationRuosong Wang, Simon S. Du, Lin F. Yang, Ruslan SalakhutdinovNeurIPS 2020 · 121 citations
