Model-Free Active Exploration in Reinforcement Learning
Alessio Russo, Alexandre Proutière
摘要
We study the problem of exploration in Reinforcement Learning and present a novel model-free solution. We adopt an information-theoretical viewpoint and start from the instance-specific lower bound of the number of samples that have to be collected to identify a nearly-optimal policy. Deriving this lower bound along with the optimal exploration strategy entails solving an intricate optimization problem and requires a model of the system. In turn, most existing sample optimal exploration algorithms rely on estimating the model. We derive an approximation of the instance-specific lower bound that only involves quantities that can be inferred using model-free approaches. Leveraging this approximation, we devise an ensemble-based model-free exploration strategy applicable to both tabular and continuous Markov decision processes. Numerical results demonstrate that our strategy is able to identify efficient policies faster than state-of-the-art exploration approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Measuring Mutual Policy Divergence for Multi-Agent Sequential ExplorationHaowen Dou, Lujuan Dang, Zhirong Luan, Badong ChenNeurIPS 2024 · 被引用 7 次
- Multi-Reward Best Policy IdentificationAlessio Russo, Filippo VannellaNeurIPS 2024 · 被引用 6 次
- In-Context Learning for Pure ExplorationAlessio Russo, Ryan Welch, Aldo PacchianoICLR 2026 · 被引用 5 次
- Adaptive Exploration for Multi-Reward Multi-Policy EvaluationAlessio Russo, Aldo PacchianoICML 2025
- Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic EnvironmentsKhang Luong, Nam Nguyen, Hoang Ta, Hung Tran-The 等ICML 2026
它引用的顶会 Paper8
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
- Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDPYuanhao Wang, Kefan Dong, Xiaoyu Chen, Liwei WangICLR 2020 · 被引用 107 次
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao 等ICML 2021 · 被引用 46 次
- Instance-Dependent Near-Optimal Policy Identification in Linear MDPs via Online Experiment DesignAndrew Wagenmaker, Kevin JamiesonNeurIPS 2022 · 被引用 38 次
相关 Paper
- Improved Sample Complexity for Reward-free Reinforcement Learning under Low-rank MDPsYuan Cheng, Ruiquan Huang, Yingbin Liang, Jing YangICLR 2023
- An Intrinsically-Motivated Approach for Learning Highly Exploring and Fast Mixing PoliciesMirco Mutti, Marcello RestelliAAAI 2020 · 被引用 31 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Provably Efficient Exploration for Reinforcement Learning Using Unsupervised LearningFei Feng, Ruosong Wang, Wotao Yin, Simon S. Du 等NeurIPS 2020 · 被引用 13 次
- Task-Optimal Exploration in Linear Dynamical SystemsAndrew J. Wagenmaker, Max Simchowitz, Kevin JamiesonICML 2021 · 被引用 24 次
