A teacher-student framework to distill future trajectories
Alexander Neitz, Giambattista Parascandolo, Bernhard Schölkopf
Abstract
By learning to predict trajectories of dynamical systems, model-based methods can make extensive use of all observations from past experience. However, due to partial observability, stochasticity, compounding errors, and irrelevant dynamics, training to predict observations explicitly often results in poor models. Model-free techniques try to side-step the problem by learning to predict values directly. While breaking the explicit dependency on future observations can result in strong performance, this usually comes at the cost of low sample efficiency, as the abundant information about the dynamics contained in future observations goes unused. Here we take a step back from both approaches: Instead of hand-designing how trajectories should be incorporated, a teacher network learns to interpret the trajectories and to provide target activations which guide a student model that can only observe the present. The teacher is trained with meta-gradients to maximize the student's performance on a validation set. We show that our approach performs well on tasks that are difficult for model-free and model-based methods, and we study the role of every component through ablation studies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c37f1c94-c4b8-4dbf-ba19-e5c87d56ed24Related papers
- Learning to Reweight Imaginary Transitions for Model-Based Reinforcement LearningWenzhen Huang, Qiyue Yin, Junge Zhang, Kaiqi HuangAAAI 2021 · 3 citations
- ODE-based Recurrent Model-free Reinforcement Learning for POMDPsXuanle Zhao, Duzhen Zhang, Liyuan Han, Tielin Zhang et al.NeurIPS 2023 · 18 citations
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 71 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing et al.NeurIPS 2020 · 12 citations
