Optimistic Exploration even with a Pessimistic Initialisation
Tabish Rashid, Bei Peng, Wendelin Boehmer, Shimon Whiteson
Abstract
Optimistic initialisation is an effective strategy for efficient exploration in reinforcement learning (RL). In the tabular case, all provably efficient model-free algorithms rely on it. However, model-free deep RL algorithms do not use optimistic initialisation despite taking inspiration from these provably efficient tabular algorithms. In particular, in scenarios with only positive rewards, Q-values are initialised at their lowest possible values due to commonly used network initialisation schemes, a pessimistic initialisation. Merely initialising the network to output optimistic Q-values is not enough, since we cannot ensure that they remain optimistic for novel state-action pairs, which is crucial for exploration. We propose a simple count-based augmentation to pessimistically initialised Q-values that separates the source of optimism from the neural network. We show that this scheme is provably efficient in the tabular setting and extend it to the deep RL setting. Our algorithm, Optimistic Pessimistically Initialised Q-Learning (OPIQ), augments the Q-value estimates of a DQN-based agent with count-derived bonuses to ensure optimism during both action selection and bootstrapping. We show that OPIQ outperforms non-optimistic DQN variants that utilise a pseudocount-based intrinsic motivation in hard exploration tasks, and that it predicts optimistic estimates for novel state-action pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99c2a699-7c52-4b28-9df8-ddbd95500df3Cited by top-tier papers13
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
- Exploit Reward Shifting in Value-Based Deep-RL: Optimistic Curiosity-Based Exploration and Conservative Exploitation via Linear Reward ShapingHao Sun, Lei Han, Rui Yang, Xiaoteng Ma et al.NeurIPS 2022 · 46 citations
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao et al.ICML 2021 · 46 citations
- Dynamic Bottleneck for Robust Self-Supervised ExplorationChenjia Bai, Lingxiao Wang, Lei Han, Animesh Garg et al.NeurIPS 2021 · 36 citations
- Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement LearningSam Lobel, Akhil Bagaria, George KonidarisICML 2023 · 29 citations
Builds on2
Related papers
- Optimistic Initialization for Exploration in Continuous ControlSam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria et al.AAAI 2022 · 14 citations
- Minimax Optimal Reinforcement Learning with Quasi-OptimismHarin Lee, Min-hwan OhICLR 2025
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- MADE: Exploration via Maximizing Deviation from Explored RegionsTianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian et al.NeurIPS 2021 · 51 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
