Active Imitation Learning with Noisy Guidance
Kianté Brantley, Hal Daumé III, Amr Sharaf
摘要
Imitation learning algorithms provide state-ofthe-art results on many structured prediction tasks by learning near-optimal search policies. Such algorithms assume training-time access to an expert that can provide the optimal action at any queried state; unfortunately, the number of such queries is often prohibitive, frequently rendering these approaches impractical. To combat this query complexity, we consider an active learning setting in which the learning algorithm has additional access to a much cheaper noisy heuristic that provides noisy guidance. Our algorithm, LEAQI, learns a difference classifier that predicts when the expert is likely to disagree with the heuristic, and queries the expert only when necessary. We apply LEAQI to three sequence labeling tasks, demonstrating significantly fewer queries to the expert and comparable (or better) accuracies over a passive approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 被引用 60 次
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter 等ICLR 2024 · 被引用 19 次
- Selective Sampling and Imitation Learning via Online RegressionAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 被引用 15 次
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter 等ICML 2023 · 被引用 13 次
- Seeing Differently, Acting Similarly: Heterogeneously Observable Imitation LearningXin-Qiang Cai, Yao-Xiang Ding, Zi-Xuan Chen, Yuan Jiang 等ICLR 2023 · 被引用 2 次
相关 Paper
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
- Teaching an Active Learner with Contrastive ExamplesChaoqi Wang, Adish Singla, Yuxin ChenNeurIPS 2021 · 被引用 17 次
- The Human-AI Substitution game: active learning from a strategic labelerTom Yan, Chicheng ZhangICLR 2024
- Efficient Active Imitation Learning with Random Network DistillationEmilien Biré, Anthony Kobanda, Ludovic Denoyer, Rémy PortelasICLR 2025
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek 等NeurIPS 2021 · 被引用 27 次
