Chain-of-Thought Predictive Control
Zhiwei Jia, Vineet Thumuluri, Fangchen Liu, Linghao Chen, Zhiao Huang, Hao Su
Abstract
We study generalizable policy learning from demonstrations for complex low-level control (e.g., contact-rich object manipulations). We propose a novel hierarchical imitation learning method that utilizes sub-optimal demos. Firstly, we propose an observation space-agnostic approach that efficiently discovers the multi-step subskill decomposition of the demos in an unsupervised manner. By grouping temporarily close and functionally similar actions into subskill-level demo segments, the observations at the segment boundaries constitute a chain of planning steps for the task, which we refer to as the chain-of-thought (CoT). Next, we propose a Transformer-based design that effectively learns to predict the CoT as the subskill-level guidance. We couple action and subskill predictions via learnable prompt tokens and a hybrid masking strategy, which enable dynamically updated guidance at test time and improve feature representation of the trajectory for generalizable policy learning. Our method, Chain-of-Thought Predictive Control (CoTPC), consistently surpasses existing strong baselines on challenging manipulation tasks with sub-optimal demos. See more details at our project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7e076a2-abf2-4d8e-b619-3386ed060f3eCited by top-tier papers8
- SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied ManipulationJunjie Zhang, Chenjia Bai, Haoran He, Zhigang Wang et al.ICML 2024 · 31 citations
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?Jingxian Lu, Wenke Xia, Yuxuan Wu, Zhiwu Lu et al.ICLR 2026 · 11 citations
- HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal DataRuizhe Liu, Pei Zhou, Qian Luo, Li Sun et al.NeurIPS 2025 · 2 citations
- Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationYuanqi Yao, Siao Liu, Haoming Song, Delin Qu et al.CVPR 2025
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningGaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel PintoICML 2025
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Chain-of-Thought Provably Enables Learning the (Otherwise) UnlearnableChenxiao Yang, Zhiyuan Li, David WipfICLR 2025
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song et al.NeurIPS 2022 · 43 citations
- Deep Bayesian Nonparametric Learning of Rules and Plans from Demonstrations with a Learned Automaton PriorBrandon Araki, Kiran Vodrahalli, Thomas Leech, Cristian Ioan Vasile et al.AAAI 2020 · 8 citations
- Generate Subgoal Images Before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal PromptsFei Ni, Jianye Hao, Shiguang Wu, Longxin Kou et al.CVPR 2024 · 4 citations
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong et al.CVPR 2026 · 43 citations
