Curriculum Design for Teaching via Demonstrations: Theory and Applications
Gaurav Yengera, Rati Devidze, Parameswaran Kamalaruban, Adish Singla
Abstract
We consider the problem of teaching via demonstrations in sequential decision-making settings. In particular, we study how to design a personalized curriculum over demonstrations to speed up the learner's convergence. We provide a unified curriculum strategy for two popular learner models: Maximum Causal Entropy Inverse Reinforcement Learning (MaxEnt-IRL) and Cross-Entropy Behavioral Cloning (CrossEnt-BC). Our unified strategy induces a ranking over demonstrations based on a notion of difficulty scores computed w.r.t. the teacher's optimal policy and the learner's current policy. Compared to the state of the art, our strategy doesn't require access to the learner's internal dynamics and still enjoys similar convergence guarantees under mild technical conditions. Furthermore, we adapt our curriculum strategy to the setting where no teacher agent is present using task-specific difficulty scores. Experiments on a synthetic car driving environment and navigation-based environments demonstrate the effectiveness of our curriculum strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 195fe9cd-5ecc-487c-acbf-4339dcfe5bb9Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira et al.ICLR 2023 · 1 citation
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Maximum Causal Entropy Specification Inference from DemonstrationsMarcell Vazquez-Chanlatte, Sanjit A. SeshiaCAV 2020 · 8 citations
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 28 citations
- Information Maximizing Curriculum: A Curriculum-Based Approach for Learning Versatile SkillsDenis Blessing, Onur Celik, Xiaogang Jia, Moritz Reuss et al.NeurIPS 2023 · 19 citations
