GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
Yangtao Chen, Zixuan Chen, Junhui Yin, Jing Huo, Pinzhuo Tian, Jieqi Shi, Yang Gao
Abstract
Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to assist in understanding novel tasks, thereby mitigating this issue. However, these methods lack a taskspecific learning process, which is essential for an accurate understanding of 3D environments, often leading to execution failures. In this paper, we introduce Grav-MAD, a sub-goal-driven, language-conditioned action diffusion framework that combines the strengths of imitation learning and foundation models. Our approach breaks tasks into sub-goals based on language instructions, allowing auxiliary guidance during both training and inference. During training, we introduce Sub-goal Keypose Discovery to identify key sub-goals from demonstrations. Inference differs from training, as there are no demonstrations available, so we use pre-trained foundation models to bridge the gap and identify sub-goals for the current task. In both phases, GravMaps are generated from sub-goals, providing GravMAD with more flexible 3D spatial guidance compared to fixed 3D positions. Empirical evaluations on RLBench show that GravMAD significantly outperforms state-of-the-art methods, with a 28.63% improvement on novel tasks and a 13.36% gain on tasks encountered during training. Evaluations on real-world robotic tasks further show that GravMAD can reason about real-world tasks, associate them with relevant visual information, and generalize to novel tasks. These results demonstrate Grav-MAD's strong multi-task learning and generalization in 3D manipulation. Video demonstrations are available at: https://gravmad.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb67ada7-cb7c-4386-8cd5-60041b240e5dCited by top-tier papers5
- Exploring the Limits of Vision-Language-Action Manipulation in Cross-task GeneralizationJiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma et al.NeurIPS 2025 · 43 citations
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action ModelsChongkai Gao, Zixuan Liu, Zhenghao Chi, Junshan Huang et al.NeurIPS 2025 · 41 citations
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D KeypointsJianshu Hu, Lidi Wang, Shujia Li, Yunpeng Jiang et al.ICLR 2026 · 6 citations
- ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon ManipulationZixuan Chen, Chongkai Gao, Lin Shao, Jieqi Shi et al.AAAI 2026 · 1 citation
- AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent PlanningZixuan Chen, Xiangrong Feng, Jieqi Shi, Lin Shao et al.CVPR 2026
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
Related papers
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingYifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang et al.AAAI 2026 · 89 citations
- Robotic Manipulation by Imitating Generated Videos Without Physical DemonstrationsShivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain et al.ICLR 2026 · 50 citations
- AutoCGP: Closed-Loop Concept-Guided Policies from Unlabeled DemonstrationsPei Zhou, Ruizhe Liu, Qian Luo, Fan Wang et al.ICLR 2025
- GenRL: Multimodal-foundation world models for generalization in embodied agentsPietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C. Courville et al.NeurIPS 2024 · 37 citations
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen et al.AAAI 2026 · 1 citation
