Pragmatically Learning from Pedagogical Demonstrations in Multi-Goal Environments
Hugo Caselles-Dupré, Olivier Sigaud, Mohamed Chetouani
摘要
Learning from demonstration methods usually leverage close to optimal demonstrations to accelerate training. By contrast, when demonstrating a task, human teachers deviate from optimal demonstrations and pedagogically modify their behavior by giving demonstrations that best disambiguate the goal they want to demonstrate. Analogously, human learners excel at pragmatically inferring the intent of the teacher, facilitating communication between the two agents. These mechanisms are critical in the few demonstrations regime, where inferring the goal is more difficult. In this paper, we implement pedagogy and pragmatism mechanisms by leveraging a Bayesian model of Goal Inference from demonstrations (BGI). We highlight the benefits of this model in multi-goal teacher-learner setups with two artificial agents that learn with goal-conditioned Reinforcement Learning. We show that combining BGI-agents (a pedagogical teacher and a pragmatic learner) results in faster learning and reduced goal ambiguity over standard learning from demonstrations, especially in the few demonstrations regime. We provide the code for our experiments (https://github.com/Caselles/NeurIPS22-demonstrations-pedagogy-pragmatism), as well as an illustrative video explaining our approach (https://youtu.be/V4n16IjkNyw).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Online Bayesian Goal Inference for Boundedly Rational Planning AgentsTan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum 等NeurIPS 2020 · 被引用 122 次
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 被引用 17 次
- Grounding Language to Autonomously-Acquired Skills via Goal GenerationAhmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani 等ICLR 2021 · 被引用 15 次
相关 Paper
- Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement LearningYantian Zha, Lin Guan, Subbarao KambhampatiAAAI 2024 · 被引用 7 次
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 被引用 25 次
- Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in ContextShashank Srivastava, Oleksandr Polozov, Nebojsa Jojic, Christopher MeekACL 2020 · 被引用 2 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
