Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement Learning
Onno Eberhard, Jakob J. Hollenstein, Cristina Pinneri, Georg Martius
Abstract
In off-policy deep reinforcement learning with continuous action spaces, exploration is often implemented by injecting action noise into the action selection process. Popular algorithms based on stochastic policies, such as SAC or MPO, inject white noise by sampling actions from uncorrelated Gaussian distributions. In many tasks, however, white noise does not provide sufficient exploration, and temporally correlated noise is used instead. A common choice is Ornstein-Uhlenbeck (OU) noise, which is closely related to Brownian motion (red noise). Both red noise and white noise belong to the broad family of colored noise. In this work, we perform a comprehensive experimental evaluation on MPO and SAC to explore the effectiveness of other colors of noise as action noise. We find that pink noise, which is halfway between white and red noise, significantly outperforms white noise, OU noise, and other alternatives on a wide range of environments. Thus, we recommend it as the default choice for action noise in continuous control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 41 citations
- Small batch deep reinforcement learningJohan S. Obando-Ceron, Marc G. Bellemare, Pablo Samuel CastroNeurIPS 2023 · 38 citations
- QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMsWei Huang, Yi Ge, Shuai Yang, Yicheng Xiao et al.ICLR 2026 · 19 citations
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow et al.ICML 2024 · 14 citations
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
Builds on1
Related papers
- Colored Noise in PPO: Improved Exploration and Performance through Correlated Action SamplingJakob J. Hollenstein, Georg Martius, Justus H. PiaterAAAI 2024 · 9 citations
- Deep Coherent Exploration for Continuous ControlYijie Zhang, Herke van HoofICML 2021 · 11 citations
- OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy EnvironmentsJinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao et al.AAAI 2024 · 14 citations
- On the Theoretical Properties of Noise Correlation in Stochastic OptimizationAurélien Lucchi, Frank Proske, Antonio Orvieto, Francis R. Bach et al.NeurIPS 2022 · 9 citations
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin et al.ICML 2024 · 8 citations
