AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
Sixiang Chen, Jiaming Liu, Siyuan Qian, Han Jiang, Zhuoyang Liu, Chenyang Gu, Xiaoqi Li, Chengkai Hou, Pengwei Wang, Zhongyuan Wang, Renrui Zhang, Shanghang Zhang
Abstract
Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicitly model the influence of the mobile base on manipulator control, which easily leads to error accumulation under high degrees of freedom. On the other hand, they treat the entire mobile manipulation process with the same visual observation modality (e.g., either all 2D or all 3D), overlooking the distinct multimodal perception requirements at different stages during mobile manipulation. To address this, we propose the Adaptive Coordination Diffusion Transformer (AC-DiT), which enhances mobile base and manipulator coordination for end-to-end mobile manipulation. First, since the motion of the mobile base directly influences the manipulator's actions, we introduce a mobility-to-body conditioning mechanism that guides the model to first extract base motion representations, which are then used as context prior for predicting whole-body actions. This enables whole-body control that accounts for the potential impact of the mobile base's motion. Second, to meet the perception requirements at different stages of mobile manipulation, we design a perceptionaware multimodal conditioning strategy that dynamically adjusts the fusion weights between various 2D visual images and 3D point clouds, yielding visual features tailored to the current perceptual needs. This allows the model to, for example, adaptively rely more on 2D inputs when semantic information is crucial for action prediction, while placing greater emphasis on 3D geometric information when precise spatial understanding is required. We empirically validate AC-DiT through extensive experiments on both simulated and real-world mobile manipulation tasks, demonstrating superior performance compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningHao Chen, Jiaming Liu, Chenyang Gu, Zhuoyang Liu et al.NeurIPS 2025 · 74 citations
- AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action ModelsXiaoqi Li, Muhe Cai, Jiadong Xu, Juan Zhu et al.CVPR 2026 · 18 citations
- Spatial Memory for Out-of-Vision Manipulation in Vision-Language-ActionPengteng Li, Weiyu Guo, He ZHANG, Tiefu Cai et al.ICML 2026 · 3 citations
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
Related papers
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksKaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang et al.AAAI 2026 · 6 citations
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang et al.ICML 2026
- SPIN: Simultaneous Perception, Interaction and NavigationShagun Uppal, Ananye Agarwal, Haoyu Xiong, Kenneth Shaw et al.CVPR 2024 · 8 citations
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan et al.ICLR 2025
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyZhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan et al.ICCV 2025 · 9 citations
