SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
Haoyi Zhu, Honghui Yang, Yating Wang, Jiange Yang, Limin Wang, Tong He
摘要
In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with intrinsic spatial understanding. We present the most comprehensive evaluation of embodied representation learning to date, covering 268 tasks across 8 simulators with diverse policies in both single-task and language-conditioned multi-task scenarios. The results are compelling: SPA consistently outperforms more than 10 state-of-the-art representation methods, including those specifically designed for embodied AI, vision-centric tasks, and multi-modal applications, while using less training data. Furthermore, we conduct a series of real-world experiments to confirm its effectiveness in practical scenarios. These results highlight the critical role of 3D spatial awareness for embodied representation learning. Our strongest model takes more than 6000 GPU hours to train and we are committed to open-sourcing all code and model weights to foster future research in embodied representation learning. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li 等NeurIPS 2025 · 被引用 79 次
- MindJourney: Test-Time Scaling with World Models for Spatial ReasoningYuncong Yang, Jiageng Liu, Zheyuan Zhang, Siyuan Zhou 等NeurIPS 2025 · 被引用 53 次
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionYixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu 等ICLR 2026 · 被引用 34 次
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual ManipulationChongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan 等CVPR 2026 · 被引用 10 次
- AETHER: Geometric-Aware Unified World ModelingHaoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang 等ICCV 2025 · 被引用 9 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic ManipulationYunlong Zhao, Xiaoheng Deng, Yichao Cao, Yi Chen 等CVPR 2026
- Abstract 3D Perception for Spatial Intelligence in Vision-Language ModelsYifan Liu, Fangneng Zhan, Kaichen Zhou, Yilun Du 等CVPR 2026 · 被引用 6 次
- Mamba-3VL: Taming State Space Model for 3D Vision Language LearningYuan Wang, Yuxin Chen, Zhongang Qi, Lijun Liu 等ICCV 2025 · 被引用 1 次
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic ScenesZhiYuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du 等ICLR 2026 · 被引用 23 次
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen 等CVPR 2026 · 被引用 64 次
