Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
Qi Sun, Pengfei Hong, Pala Tej Deep, Vernon Toh, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria
Abstract
Traditional reinforcement learning-based robotic control methods are often taskspecific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, EMMA-X. EMMA-X leverages our constructed hierarchical embodiment * Both authors contributed equally to this work. The first authorship was randomly assigned by coin flip. † Now at Deepmind. dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that EMMA-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning. We make our codes, models and datasets publicly available: https: //declare-lab.github.io/Emma-X/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffc1c1e8-5459-4442-aea4-c8f60b8c2f1aCited by top-tier papers10
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
- Vlaser: Vision-Language-Action Model with Synergistic Embodied ReasoningGanlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang et al.ICLR 2026 · 23 citations
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual ForesightYi Yang, Xueqi Li, Yiyang Chen, Jin Song et al.CVPR 2026 · 14 citations
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?Jingxian Lu, Wenke Xia, Yuxuan Wu, Zhiwu Lu et al.ICLR 2026 · 11 citations
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li et al.CVPR 2026 · 11 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage et al.EMNLP 2023 · 1 citation
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong et al.CVPR 2026
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang et al.NeurIPS 2025 · 179 citations
- E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial IntelligenceJunbo Qi, Yi Zhang, Hanchu Ni, Che Liu et al.ACL 2026
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang et al.ICML 2024 · 303 citations
