Student-Informed Teacher Training
Nico Messikommer, Jiaxu Xing, Elie Aljalbout, Davide Scaramuzza
Abstract
Imitation learning with a privileged teacher has proven effective for learning complex control behaviors from high-dimensional inputs, such as images. In this framework, a teacher is trained with privileged task information, while a student tries to predict the actions of the teacher with more limited observations, e.g., in a robot navigation task, the teacher might have access to distances to nearby obstacles, while the student only receives visual observations of the scene. However, privileged imitation learning faces a key challenge: the student might be unable to imitate the teacher's behavior due to partial observability. This problem arises because the teacher is trained without considering if the student is capable of imitating the learned behavior. To address this teacher-student asymmetry, we propose a framework for joint training of the teacher and student policies, encouraging the teacher to learn behaviors that can be imitated by the student despite the latters' limited access to information and its partial observability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End DrivingLong Nguyen, Micha Fauth, Bernhard Jaeger, Daniel Dauner et al.CVPR 2026 · 28 citations
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
- Uncertainty-Sensitive Privileged LearningFan-Ming Luo, Lei Yuan, Yang YuNeurIPS 2025
Builds on1
Related papers
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Robust Asymmetric Learning in POMDPsAndrew Warrington, Jonathan Wilder Lavington, Adam Scibior, Mark Schmidt et al.ICML 2021 · 37 citations
- Coaching a Teachable StudentJimuyang Zhang, Zanming Huang, Eshed Ohn-BarCVPR 2023
- Real-World Reinforcement Learning of Active Perception BehaviorsEdward S. Hu, Jie Wang, Xingfang Yuan, Fiona Luo et al.NeurIPS 2025 · 6 citations
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 39 citations
