Embodied CoT Distillation From LLM To Off-the-shelf Agents
Wonje Choi, Woo Kyung Kim, Minjong Yoo, Honguk Woo
Abstract
We address the challenge of utilizing large language models (LLMs) for complex embodied tasks, in the environment where decision-making systems operate timely on capacity-limited, off-the-shelf devices. We present DeDer, a framework for decomposing and distilling the embodied reasoning capabilities from LLMs to efficient, small language model (sLM)-based policies. In DeDer, the decision-making process of LLM-based strategies is restructured into a hierarchy with a reasoning-policy and planning-policy. The reasoning-policy is distilled from the data that is generated through the embodied in-context learning and self-verification of an LLM, so it can produce effective rationales. The planning-policy, guided by the rationales, can render optimized plans efficiently. In turn, DeDer allows for adopting sLMs for both policies, deployed on off-the-shelf devices. Furthermore, to enhance the quality of intermediate rationales, specific to embodied tasks, we devise the embodied knowledge graph, and to generate multiple rationales timely through a single inference, we also use the contrastively prompted attention model. Our experiments with the ALFRED benchmark demonstrate that DeDer surpasses leading language planning and distillation approaches, indicating the applicability and efficiency of sLM-based embodied policies derived through DeDer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f7f8a8c-d90f-4aeb-bd21-5fe4d55c706dCited by top-tier papers5
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
- TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic EnvironmentsZhiyu Huang, Yun Zhang, Johnson Liu, Rui Song et al.ICML 2026 · 11 citations
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied ReasoningWonje Choi, Jooyoung Kim, Honguk WooNeurIPS 2025 · 4 citations
- Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool UsageZhi Gao, Bofei Zhang, Pengxiang Li, Xiaojian Ma et al.ICLR 2025
- Efficient Skill Grounding via Code Refactoring with Small Language ModelsSera Choi, Wonje Choi, Saehun Chun, Daehee Lee et al.ICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction FollowingMinjong Yoo, Jinwoo Jang, Wei-Jin Park, Honguk WooNeurIPS 2024 · 15 citations
- PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningFaeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang et al.ICLR 2024 · 8 citations
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model ReasoningLi Wang, Changhao Zhang, Zengqi Xiu, Kai Lu et al.AAAI 2026 · 1 citation
- Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive TasksMinki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi et al.NeurIPS 2023 · 128 citations
- LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied AgentsJae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim et al.ICLR 2024 · 49 citations
