Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
Qi Sun, Pengfei Hong, Pala Tej Deep, Vernon Toh, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria
摘要
Traditional reinforcement learning-based robotic control methods are often taskspecific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, EMMA-X. EMMA-X leverages our constructed hierarchical embodiment * Both authors contributed equally to this work. The first authorship was randomly assigned by coin flip. † Now at Deepmind. dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that EMMA-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning. We make our codes, models and datasets publicly available: https: //declare-lab.github.io/Emma-X/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen 等ICLR 2026 · 被引用 50 次
- Vlaser: Vision-Language-Action Model with Synergistic Embodied ReasoningGanlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang 等ICLR 2026 · 被引用 23 次
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual ForesightYi Yang, Xueqi Li, Yiyang Chen, Jin Song 等CVPR 2026 · 被引用 14 次
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?Jingxian Lu, Wenke Xia, Yuxuan Wu, Zhiwu Lu 等ICLR 2026 · 被引用 11 次
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li 等CVPR 2026 · 被引用 11 次
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage 等EMNLP 2023 · 被引用 1 次
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong 等CVPR 2026
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang 等NeurIPS 2025 · 被引用 179 次
- E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial IntelligenceJunbo Qi, Yi Zhang, Hanchu Ni, Che Liu 等ACL 2026
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 等ICML 2024 · 被引用 303 次
