SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
Zekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong, Xinqiang Yu, Jingwen Li, Lingyun Xu, Baoyu Li, Xialin He, Guofan Fan, Jiazhao Zhang, Jiawei He
Abstract
While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation-a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the"plug-in"direction of a USB or the"handle"direction of a cup). To support this, we construct OrienText300K, a large-scale dataset of 3D objects annotated with semantic orientations, and develop PointSO, a general model for zero-shot semantic orientation prediction. By integrating semantic orientation into VLM agents, our SoFar framework enables 6-DoF spatial reasoning and generates robotic actions. Extensive experiments demonstrated the effectiveness and generalization of our SoFar, e.g., zero-shot 48.7% successful rate on Open6DOR and zero-shot 74.9% successful rate on SIMPLER-Env.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd56328f-dc76-49ad-b748-aa071f2dc1d5Cited by top-tier papers24
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang et al.ICLR 2026 · 109 citations
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic ManipulationYifu Yuan, Haiqin Cui, Yaoting Huang, Yibin Chen et al.ICLR 2026 · 48 citations
- Orient Anything V2: Unifying Orientation and Rotation UnderstandingZehan Wang, Ziang Zhang, Jiayang Xu, Jialei Wang et al.NeurIPS 2025 · 24 citations
- Diagnose, Correct, and Learn from Manipulation Failures via Visual SymbolsXianchao Zeng, Xinyu Zhou, Youcheng Li, Jiayou Shi et al.CVPR 2026 · 18 citations
Builds on62
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- BOP-ASK: Object-Interaction Reasoning for Vision-Language ModelsVineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich et al.CVPR 2026 · 6 citations
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept VectorsLiming Kuang, Yordanka Velikova, Mahdi Saleh, Jan-Nico Zaech et al.CVPR 2026 · 5 citations
- GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language InstructionsXiaomeng Chu, Jiajun Deng, Guoliang You, Wei Liu et al.ICCV 2025
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic ManipulationYifu Yuan, Haiqin Cui, Yibin Chen, Zibin Dong et al.ICLR 2026 · 41 citations
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
