Lune

CVPR2025顶会

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, Stan Birchfield

2025年份
50顶会引用

摘要

Generalist robot policies require strong spatial priors to operate reliably across diverse environments, enabling them to perceive, reason, and act within 3D space from multiple perspectives. Vision-language models (VLMs) are promising backbones for such policies but are limited by training on generic web-scale image-text datasets that lack rich, multi-frame spatial cues for manipulation. One example is reference frame comprehension-deciding whether to reason in egocentric, world-centric, or object-centric coordinates-which is critical for precise, context-aware actions. We introduce ROBOSPATIAL, a large-scale dataset built from real indoor and tabletop 3D scans paired with egocentric RGB views, containing 1M images, 5k scans, and 3M annotated spatial relations spanning objectobject, object-space, and object-compatibility reasoning. Its 2D/3D-ready design supports learning priors that generalize across viewpoints, scales, and task contexts. Models trained on ROBOSPATIAL achieve significant gains in spatial reasoning benchmarks and robot manipulation, demonstrating how targeted spatial priors enhance the generalization and reliability of robot policies.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper50

问问它们各自怎么用它

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖