Efficient Active Imitation Learning with Random Network Distillation
Emilien Biré, Anthony Kobanda, Ludovic Denoyer, Rémy Portelas
摘要
Developing agents for complex and underspecified tasks, where no clear objective exists, remains challenging but offers many opportunities. This is especially true in video games, where simulated players (bots) need to play realistically, and there is no clear reward to evaluate them. While imitation learning has shown promise in such domains, these methods often fail when agents encounter out-of-distribution scenarios during deployment. Expanding the training dataset is a common solution, but it becomes impractical or costly when relying on human demonstrations. This article addresses active imitation learning, aiming to trigger expert intervention only when necessary, reducing the need for constant expert input along training. We introduce Random Network Distillation DAgger (RND-DAgger), a new active imitation learning method that limits expert querying by using a learned state-based out-of-distribution measure to trigger interventions. This approach avoids frequent expert-agent action comparisons, thus making the expert intervene only when it is useful. We evaluate RND-DAgger against traditional imitation learning and other active approaches in 3D video games (racing and third-person navigation) and in a robotic locomotion task and show that RND-DAgger surpasses previous methods by reducing expert queries. https://sites.google.com/view/rnd-dagger
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EAPO: Enhancing Policy Optimization with On-Demand Expert AssistanceSiyao Song, Cong Ma, Zhihao Cheng, Shiye Lei 等ICML 2026
- Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret MinimizationBo Ling, Zhengyu Gan, Wanyuan Wang, Guanyu Gao 等NeurIPS 2025
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention MechanismHaoyuan Cai, Zhenghao Peng, Bolei ZhouICML 2025
- Adaptive Quasimetric Mapping : Principled Topological Abstraction for Robust Offline Goal-Conditioned NavigationAnthony Kobanda, Waris Radji, Odalric-Ambrym Maillard, Rémy PortelasICML 2026
它引用的顶会 Paper2
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- Stylized Offline Reinforcement Learning: Extracting Diverse High-Quality Behaviors from Heterogeneous DatasetsYihuan Mao, Chengjie Wu, Xi Chen, Hao Hu 等ICLR 2024 · 被引用 9 次
相关 Paper
- Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent FeedbackMichelle D. Zhao, Henny Admoni, Reid G. Simmons, Aaditya Ramdas 等ICLR 2025
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
- Interactive and Hybrid Imitation Learning: Provably Beating Behavior CloningYichen Li, Chicheng ZhangNeurIPS 2025 · 被引用 1 次
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
- Active Imitation Learning with Noisy GuidanceKianté Brantley, Hal Daumé III, Amr SharafACL 2020 · 被引用 2 次
