Deep Coherent Exploration for Continuous Control
Yijie Zhang, Herke van Hoof
摘要
In policy search methods for reinforcement learning (RL), exploration is often performed by injecting noise either in action space at each step independently or in parameter space over each full trajectory. In prior work, it has been shown that with linear policies, a more balanced trade-off between these two exploration strategies is beneficial. However, that method did not scale to policies using deep neural networks. In this paper, we introduce deep coherent exploration, a general and scalable exploration framework for deep RL algorithms for continuous control, that generalizes step-based and trajectory-based exploration. This framework models the last layer parameters of the policy network as latent variables and uses a recursive inference step within the policy update to handle these latent variables in a scalable manner. We find that deep coherent exploration improves the speed and stability of learning of A2C, PPO, and SAC on several continuous control tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin 等ICML 2024 · 被引用 8 次
- From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous EnvironmentsSaket Tiwari, Tejas Kotwal, George Dimitri KonidarisICLR 2026
- Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement LearningOnno Eberhard, Jakob J. Hollenstein, Cristina Pinneri, Georg MartiusICLR 2023
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 被引用 41 次
