Chain of Thought Imitation with Procedure Cloning
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, Ofir Nachum
摘要
Imitation learning aims to extract high-performance policies from logged demonstrations of expert behavior. It is common to frame imitation learning as a supervised learning problem in which one fits a function approximator to the input-output mapping exhibited by the logged demonstrations (input observations to output actions). While the framing of imitation learning as a supervised input-output learning problem allows for applicability in a wide variety of settings, it is also an overly simplistic view of the problem in situations where the expert demonstrations provide much richer insight into expert behavior. For example, applications such as path navigation, robot manipulation, and strategy games acquire expert demonstrations via planning, search, or some other multi-step algorithm, revealing not just the output action to be imitated but also the procedure for how to determine this action. While these intermediate computations may use tools not available to the agent during inference (e.g., environment simulators), they are nevertheless informative as a way to explain an expert's mapping of state to actions. To properly leverage expert procedure information without relying on the privileged tools the expert may have used to perform the procedure, we propose procedure cloning, which applies supervised sequence prediction to imitate the series of expert computations. This way, procedure cloning learns not only what to do (i.e., the output action), but how and why to do it (i.e., the procedure). Through empirical analysis on navigation, simulated robotic manipulation, and game-playing environments, we show that imitating the intermediate computations of an expert's behavior enables procedure cloning to learn policies exhibiting significant generalization to unseen environment configurations, including those configurations for which running the expert's procedure directly is infeasible. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 被引用 539 次
- Chain-of-Thought Predictive ControlZhiwei Jia, Vineet Thumuluri, Fangchen Liu, Linghao Chen 等ICML 2024 · 被引用 24 次
- CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual GroundingEslam Mohamed Bakr, Mohamed Ayman, Mahmoud Ahmed, Habib Slim 等ICLR 2024 · 被引用 16 次
- Multi-Environment Pretraining Enables Transfer to Action Limited DatasetsDavid Venuto, Sherry Yang, Pieter Abbeel, Doina Precup 等ICML 2023 · 被引用 7 次
- Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language ModelsXiao-Wen Yang, Zi-Yu Han, Xi-Hua Zhang, Wen-Da Wei 等ICML 2026 · 被引用 6 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
相关 Paper
- When a Robot is More Capable than a Human: Learning from Constrained DemonstratorsXinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz 等ICLR 2026
- Diffusion Model-Augmented Behavioral CloningShang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai 等ICML 2024 · 被引用 47 次
- Provable Representation Learning for Imitation Learning via Bi-level OptimizationSanjeev Arora, Simon S. Du, Sham M. Kakade, Yuping Luo 等ICML 2020 · 被引用 65 次
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette 等ICLR 2023 · 被引用 2 次
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 被引用 26 次
