Interactive Learning from Activity Description
Khanh Nguyen, Dipendra Misra, Robert E. Schapire, Miroslav Dudík, Patrick Shafto
摘要
We present a novel interactive learning protocol that enables training request-fulfilling agents by verbally describing their activities. Unlike imitation learning (IL), our protocol allows the teaching agent to provide feedback in a language that is most appropriate for them. Compared with reward in reinforcement learning (RL), the description feedback is richer and allows for improved sample complexity. We develop a probabilistic framework and an algorithm that practically implements our protocol. Empirical results in two challenging request-fulfilling problems demonstrate the strengths of our approach: compared with RL baselines, it is more sample-efficient; compared with IL baselines, it achieves competitive success rates without requiring the teaching agent to be able to demonstrate the desired behavior using the learning agent's actions. Apart from empirical evaluation, we also provide theoretical guarantees for our algorithm under certain assumptions about the teacher and the environment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro 等NeurIPS 2024 · 被引用 102 次
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- Passive learning of active causal strategies in agents and language modelsAndrew K. Lampinen, Stephanie C. Y. Chan, Ishita Dasgupta, Andrew J. Nam 等NeurIPS 2023 · 被引用 30 次
- Imitating Past Successes can be Very SuboptimalBenjamin Eysenbach, Soumith Udatha, Russ Salakhutdinov, Sergey LevineNeurIPS 2022 · 被引用 28 次
它引用的顶会 Paper5
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 被引用 271 次
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 等NeurIPS 2020 · 被引用 139 次
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan 等AAAI 2021 · 被引用 67 次
- Inverse Reinforcement Learning with Natural Language GoalsLi Zhou, Kevin SmallAAAI 2021 · 被引用 40 次
- An Imitation Game for Learning Semantic Parsers from User InteractionZiyu Yao, Yiqi Tang, Wen-tau Yih, Huan Sun 等EMNLP 2020 · 被引用 18 次
相关 Paper
- Learning to Interactively Learn and AssistMark Woodward, Chelsea Finn, Karol HausmanAAAI 2020 · 被引用 37 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Imitation Learning as Return Distribution MatchingFilippo Lazzati, Alberto Maria MetelliICLR 2026 · 被引用 1 次
- Policy Improvement using Language Feedback ModelsVictor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre CôtéNeurIPS 2024 · 被引用 18 次
- Provably Feedback-Efficient Reinforcement Learning via Active Reward LearningDingwen Kong, Lin YangNeurIPS 2022 · 被引用 19 次
