Latent exploration for Reinforcement Learning
Alberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander Mathis
Abstract
In Reinforcement Learning, agents learn policies by exploring and interacting with the environment. Due to the curse of dimensionality, learning policies that map high-dimensional sensory input to motor output is particularly challenging. During training, state of the art methods (SAC, PPO, etc.) explore the environment by perturbing the actuation with independent Gaussian noise. While this unstructured exploration has proven successful in numerous tasks, it can be suboptimal for overactuated systems. When multiple actuators, such as motors or muscles, drive behavior, uncorrelated perturbations risk diminishing each other's effect, or modifying the behavior in a task-irrelevant way. While solutions to introduce time correlation across action perturbations exist, introducing correlation across actuators has been largely ignored. Here, we propose LATent TIme-Correlated Exploration (Lattice), a method to inject temporally-correlated noise into the latent state of the policy network, which can be seamlessly integrated with on- and off-policy algorithms. We demonstrate that the noisy actions generated by perturbing the network's activations can be modeled as a multivariate Gaussian distribution with a full covariance matrix. In the PyBullet locomotion tasks, Lattice-SAC achieves state of the art results, and reaches 18% higher reward than unstructured exploration in the Humanoid environment. In the musculoskeletal control environments of MyoSuite, Lattice-PPO achieves higher reward in most reaching and object manipulation tasks, while also finding more energy-efficient policies with reductions of 20-60%. Overall, we demonstrate the effectiveness of structured action noise in time and actuator space for complex motor control tasks. The code is available at: https://github.com/amathislab/lattice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75b206b9-8a4e-4510-bb00-d145c40759fdCited by top-tier papers12
- DynSyn: Dynamical Synergistic Representation for Efficient Learning and Control in Overactuated Embodied SystemsKaibo He, Chenhui Zuo, Chengtian Ma, Yanan SuiICML 2024 · 19 citations
- Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement LearningGe Li, Hongyi Zhou, Dominik Roth, Serge Thilges et al.ICLR 2024 · 11 citations
- Colored Noise in PPO: Improved Exploration and Performance through Correlated Action SamplingJakob J. Hollenstein, Georg Martius, Justus H. PiaterAAAI 2024 · 9 citations
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 8 citations
- Demystifying Reward Design in Reinforcement Learning for Upper Extremity Interaction: Practical Guidelines for Biomechanical Simulations in HCIHannah Selder, Florian Fischer, Per Ola Kristensson, Arthur FleigUIST 2025 · 2 citations
Builds on3
- DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing bodyAlberto Silvio Chiappa, Alessandro Marin Vargas, Alexander MathisNeurIPS 2022 · 13 citations
- DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal SystemsPierre Schumacher, Daniel F. B. Haeufle, Dieter Büchler, Syn Schmitt et al.ICLR 2023 · 6 citations
- Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement LearningOnno Eberhard, Jakob J. Hollenstein, Cristina Pinneri, Georg MartiusICLR 2023
Related papers
- Deep Coherent Exploration for Continuous ControlYijie Zhang, Herke van HoofICML 2021 · 11 citations
- Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated ControlYiming Wang, Kaiyan Zhao, Xu Li, Yan Li et al.AAAI 2026 · 1 citation
- Latent State-Predictive Exploration for Deep Reinforcement LearningYiming Wang, Kaiyan Zhao, Borong Zhang, Yan Li et al.AAAI 2026 · 1 citation
- Recomposing the Reinforcement Learning Building Blocks with HypernetworksElad Sarafian, Shai Keynan, Sarit KrausICML 2021 · 42 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
