Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
Ev Zisselman, Mirco Mutti, Shelly Francis-Meretzki, Elisei Shafer, Aviv Tamar
摘要
Behavioral cloning is a simple yet effective technique for learning sequential decision-making from demonstrations. Recently, it has gained prominence as the core of foundation models for the physical world, where achieving generalization requires countless demonstrations of a multitude of tasks. Typically, a human expert with full information on the task demonstrates a (nearly) optimal behavior. In this paper, we propose to hide some of the task's information from the demonstrator. This ``blindfolded''expert is compelled to employ non-trivial exploration to solve the task. We show that cloning the blindfolded expert generalizes better to unseen tasks than its fully-informed counterpart. We conduct experiments of real-world robot peg insertion tasks with (limited) human demonstrations, alongside videogames from the Procgen benchmark. Additionally, we support our findings with theoretical analysis, which confirms that the generalization error scales with , where measures the amount of task information available to the demonstrator, and is the number of demonstrated tasks. Both theory and practice indicate that cloning blindfolded experts generalizes better with fewer demonstrated tasks. Project page with videos and code: https://sites.google.com/view/blindfoldedexperts/home
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
- Understanding Domain Randomization for Sim-to-real TransferXiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li 等ICLR 2022 · 被引用 164 次
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov 等NeurIPS 2021 · 被引用 143 次
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 被引用 141 次
相关 Paper
- When a Robot is More Capable than a Human: Learning from Constrained DemonstratorsXinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz 等ICLR 2026
- Chain of Thought Imitation with Procedure CloningMengjiao Yang, Dale Schuurmans, Pieter Abbeel, Ofir NachumNeurIPS 2022 · 被引用 53 次
- The Generalization Gap in Offline Reinforcement LearningIshita Mediratta, Qingfei You, Minqi Jiang, Roberta RaileanuICLR 2024 · 被引用 24 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj 等NeurIPS 2025 · 被引用 27 次
