Teachable Reinforcement Learning via Advice Distillation
Olivia Watkins, Abhishek Gupta, Trevor Darrell, Pieter Abbeel, Jacob Andreas
摘要
Training automated agents to complete complex tasks in interactive environments is challenging: reinforcement learning requires careful hand-engineering of reward functions, imitation learning requires specialized infrastructure and access to a human expert, and learning from intermediate forms of supervision (like binary preferences) is time-consuming and extracts little information from each human intervention. Can we overcome these challenges by building agents that learn from rich, interactive feedback instead? We propose a new supervision paradigm for interactive learning based on "teachable" decision-making systems that learn from structured advice provided by an external teacher. We begin by formalizing a class of human-in-the-loop decision making problems in which multiple forms of teacher-provided advice are available to a learner. We then describe a simple learning algorithm for these problems that first learns to interpret advice, then learns from advice to complete tasks even in the absence of human supervision. In puzzle-solving, navigation, and locomotion domains, we show that agents that learn from advice can acquire new skills with significantly less human supervision than standard reinforcement learning algorithms and often less than imitation learning. 35th Conference on Neural Information Processing Systems (NeurIPS 2021), virtual.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Hierarchically Decoupled Imitation For Morphological TransferDonald J. Hejna III, Lerrel Pinto, Pieter AbbeelICML 2020 · 被引用 47 次
- Interactive Learning from Activity DescriptionKhanh Nguyen, Dipendra Misra, Robert E. Schapire, Miroslav Dudík 等ICML 2021 · 被引用 36 次
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
相关 Paper
- A Framework for Learning to Request Rich and Contextually Useful Information from HumansKhanh X. Nguyen, Yonatan Bisk, Hal Daumé IIIICML 2022 · 被引用 21 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- Provable Interactive Learning with Hindsight Instruction FeedbackDipendra Misra, Aldo Pacchiano, Robert E. SchapireICML 2024 · 被引用 1 次
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationLin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang 等NeurIPS 2021 · 被引用 57 次
- Provably Feedback-Efficient Reinforcement Learning via Active Reward LearningDingwen Kong, Lin YangNeurIPS 2022 · 被引用 19 次
