Pragmatically Learning from Pedagogical Demonstrations in Multi-Goal Environments
Hugo Caselles-Dupré, Olivier Sigaud, Mohamed Chetouani
Abstract
Learning from demonstration methods usually leverage close to optimal demonstrations to accelerate training. By contrast, when demonstrating a task, human teachers deviate from optimal demonstrations and pedagogically modify their behavior by giving demonstrations that best disambiguate the goal they want to demonstrate. Analogously, human learners excel at pragmatically inferring the intent of the teacher, facilitating communication between the two agents. These mechanisms are critical in the few demonstrations regime, where inferring the goal is more difficult. In this paper, we implement pedagogy and pragmatism mechanisms by leveraging a Bayesian model of Goal Inference from demonstrations (BGI). We highlight the benefits of this model in multi-goal teacher-learner setups with two artificial agents that learn with goal-conditioned Reinforcement Learning. We show that combining BGI-agents (a pedagogical teacher and a pragmatic learner) results in faster learning and reduced goal ambiguity over standard learning from demonstrations, especially in the few demonstrations regime. We provide the code for our experiments (https://github.com/Caselles/NeurIPS22-demonstrations-pedagogy-pragmatism), as well as an illustrative video explaining our approach (https://youtu.be/V4n16IjkNyw).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f86c917-9d44-43bc-b363-af6c8d224b8eBuilds on4
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Online Bayesian Goal Inference for Boundedly Rational Planning AgentsTan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum et al.NeurIPS 2020 · 122 citations
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 17 citations
- Grounding Language to Autonomously-Acquired Skills via Goal GenerationAhmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani et al.ICLR 2021 · 15 citations
Related papers
- Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement LearningYantian Zha, Lin Guan, Subbarao KambhampatiAAAI 2024 · 7 citations
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin et al.ICML 2022 · 32 citations
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 25 citations
- Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in ContextShashank Srivastava, Oleksandr Polozov, Nebojsa Jojic, Christopher MeekACL 2020 · 2 citations
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 22 citations
