Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
Edwin Zhang, Yujie Lu, Shinda Huang, William Yang Wang, Amy Zhang
Abstract
Training generalist agents is difficult across several axes, requiring us to deal with high-dimensional inputs (space), long horizons (time), and generalization to novel tasks. Recent advances with architectures have allowed for improved scaling along one or two of these axes, but are still computationally prohibitive to use. In this paper, we propose to address all three axes by leveraging Language to Control Diffusion models as a hierarchical planner conditioned on language (LCD). We effectively and efficiently scale diffusion models for planning in extended temporal, state, and task dimensions to tackle long horizon control problems conditioned on natural language instructions, as a step towards generalist agents. Comparing LCD with other state-of-the-art models on the CALVIN language robotics benchmark finds that LCD outperforms other SOTA methods in multi-task success rates, whilst improving inference speed over other comparable diffusion models by 3.3x 15x. We show that LCD can successfully leverage the unique strength of diffusion models to produce coherent long range plans while addressing their weakness in generating low-level details and control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han et al.NeurIPS 2024 · 153 citations
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyZhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan et al.ICCV 2025 · 9 citations
- LLM-based Skill Diffusion for Zero-shot Policy AdaptationWoo Kyung Kim, Youngseok Lee, Jooyoung Kim, Honguk WooNeurIPS 2024 · 7 citations
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action PolicyTianyi Zhang, Haonan Duan, Haoran Hao, Yu Qiao et al.AAAI 2026 · 5 citations
- RoboTron-Mani: All-in-One Multimodal Large Model for Robotic ManipulationFeng Yan, Fanfan Liu, Yiyang Huang, Zechao Guan et al.ICCV 2025 · 1 citation
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Zero-Shot Robotic Manipulation with Pre-Trained Image-Editing Diffusion ModelsKevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke et al.ICLR 2024 · 284 citations
- SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task ExecutionZhixuan Liang, Yao Mu, Hengbo Ma, Masayoshi Tomizuka et al.CVPR 2024
- Pixel Motion Diffusion is What We Need for Robot ControlE-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li et al.CVPR 2026 · 10 citations
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng et al.ICLR 2026 · 18 citations
- Extendable Planning via Multiscale DiffusionChang Chen, Hany Hamed, Doojin Baek, Taegu Kang et al.AAAI 2026 · 3 citations
