OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy Environments
Jinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao, Chenjia Bai, Junjie Ye, Zhen Wang, Haiyin Piao, Yang Sun
Abstract
In reinforcement learning, the optimism in the face of uncertainty (OFU) is a mainstream principle for directing exploration towards less explored areas, characterized by higher uncertainty. However, in the presence of environmental stochasticity (noise), purely optimistic exploration may lead to excessive probing of high-noise areas, consequently impeding exploration efficiency. Hence, in exploring noisy environments, while optimism-driven exploration serves as a foundation, prudent attention to alleviating unnecessary over-exploration in high-noise areas becomes beneficial. In this work, we propose Optimistic Value Distribution Explorer (OVD-Explorer) to achieve a noise-aware optimistic exploration for continuous control. OVD-Explorer proposes a new measurement of the policy's exploration ability considering noise in optimistic perspectives, and leverages gradient ascent to drive exploration. Practically, OVD-Explorer can be easily integrated with continuous control RL algorithms. Extensive evaluations on the MuJoCo and GridChaos tasks demonstrate the superiority of OVD-Explorer in achieving noise-aware optimistic exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2e49fd6-b187-4b32-8428-4f2910c98edeCited by top-tier papers3
- EvoRainbow: Combining Improvements in Evolutionary Reinforcement Learning for Policy SearchPengyi Li, Yan Zheng, Hongyao Tang, Xian Fu et al.ICML 2024 · 13 citations
- Value-Evolutionary-Based Reinforcement LearningPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng et al.ICML 2024 · 10 citations
- Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement LearningYiwen Zhu, Jinyi Liu, Pengjie Gu, Yifu Yuan et al.NeurIPS 2025 · 1 citation
Builds on12
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 86 citations
- HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action RepresentationBoyan Li, Hongyao Tang, Yan Zheng, Jianye Hao et al.ICLR 2022 · 79 citations
Related papers
- Efficient Model-Based Reinforcement Learning Through Optimistic Thompson SamplingJasmine Bayrooti, Carl Henrik Ek, Amanda ProrokICLR 2025
- Bayesian Optimistic Optimization: Optimistic Exploration for Model-based Reinforcement LearningChenyang Wu, Tianci Li, Zongzhang Zhang, Yang YuNeurIPS 2022 · 9 citations
- SHAPO: Sharpness-Aware Policy Optimization for Safe ExplorationKaustubh Mani, Yann Pequignot, Vincent Mai, Liam PaullICLR 2026 · 3 citations
- Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk CriterionTaehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee et al.NeurIPS 2023 · 10 citations
- Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement LearningOnno Eberhard, Jakob J. Hollenstein, Cristina Pinneri, Georg MartiusICLR 2023
