Data augmentation for efficient learning from parametric experts
Alexandre Galashov, Joshua Scott Merel, Nicolas Heess
Abstract
We present a simple, yet powerful data-augmentation technique to enable data-efficient learning from parametric experts for reinforcement and imitation learning. We focus on what we call the policy cloning setting, in which we use online or offline queries of an expert or expert policy to inform the behavior of a student policy. This setting arises naturally in a number of problems, for instance as variants of behavior cloning, or as a component of other algorithms such as DAGGER, policy distillation or KL-regularized RL. Our approach, augmented policy cloning (APC), uses synthetic states to induce feedback-sensitivity in a region around sampled trajectories, thus dramatically reducing the environment interactions required for successful cloning of the expert. We achieve highly data-efficient transfer of behavior from an expert to a student policy for high-degrees-of-freedom control problems. We demonstrate the benefit of our method in the context of several existing and widely used algorithms that include policy cloning as a constituent part. Moreover, we highlight the benefits of our approach in two practically relevant settings (a) expert compression, i.e. transfer to a student with fewer parameters; and (b) transfer from privileged experts, i.e. where the expert has a different observation space than the student, usually including access to privileged information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8fa4105-06c8-48a3-a80f-2fa92fdf5219Cited by top-tier papers2
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 135 citations
- Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a SimulatorJunmin Zhong, Emiliano Quiñones Yumbla, Seyed Yousef Soltanian, Ruofan Wu et al.ICML 2025
Builds on5
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark et al.ICLR 2020 · 138 citations
- Efficient Transformers in Reinforcement Learning using Actor-Learner DistillationEmilio Parisotto, Ruslan SalakhutdinovICLR 2021 · 51 citations
- RL Unplugged: A Collection of Benchmarks for Offline Reinforcement LearningÇaglar Gülçehre, Ziyu Wang, Alexander Novikov, Thomas Paine et al.NeurIPS 2020 · 25 citations
Related papers
- TaSIL: Taylor Series Imitation LearningDaniel Pfrommer, Thomas T. C. K. Zhang, Stephen Tu, Nikolai MatniNeurIPS 2022 · 27 citations
- Demonstration-Regularized RLDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.ICLR 2024 · 5 citations
- Difference-Aware Retrieval Policies for Imitation LearningQuinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal et al.ICLR 2026 · 1 citation
- Provable Guarantees for Generative Behavior Cloning: Bridging Low-Level Stability and High-Level BehaviorAdam Block, Ali Jadbabaie, Daniel Pfrommer, Max Simchowitz et al.NeurIPS 2023 · 44 citations
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
