SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang
Abstract
The window to the right of the oven hood." "The laptop beside the floral-patterned chair." "The bookshelf second from the right. " "The chair backed to the window." "The closed door. Not the bathroom door. " "Select the couch that has an L shape." Previous SoTA Ours (a) Texture (b) Shape (c) Viewpoint (d) Orientation (e) State (f) Order Closed Open Right Left Right Illustration Flor. chair Chair back L shape I shape Window Window Lap. 1 st 3 rd 4 th 2 nd Window Oven hood Figure 1. Effectiveness of SeeGround: Different from previous SoTA, our method associates 2D visual cues -color, texture, viewpoint, spatial position, orientation, and state -with 3D spatial text description to achieve precise scene understanding. Specifically, our method: (a) identifies the floral chair by recognizing unique color and texture cues; (b) recognizes the couch by interpreting geometric shape; (c) determines the right window by interpreting spatial relationships and perspective; (d) identifies the chair by discerning directional alignment; (e) detects the closed door by visually interpreting its state; and (f) selects the bookshelf by understanding relative positioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 815cb8d6-ab37-4065-9813-46d1fdc75253Cited by top-tier papers19
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong et al.ICLR 2026 · 28 citations
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual GroundingZhao Jin, Rong-Cheng Tu, Jingyi Liao, Wenhao Sun et al.NeurIPS 2025 · 13 citations
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR SequencesXiang Xu, Ao Liang, Youquan Liu, Linfeng Li et al.CVPR 2026 · 8 citations
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
- 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive DecodingMakanjuola Adekunmi Ogunleye, Eman Abdelrahman, Ismini LourentzouCVPR 2026 · 3 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
Related papers
- OpenScene: 3D Scene Understanding with Open VocabulariesSongyou Peng, Kyle Genova, Chiyu Max Jiang, Andrea Tagliasacchi et al.CVPR 2023
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingHongjia Zhai, Hai Li, Zhenzhe Li, Xiaokun Pan et al.CVPR 2025
- Text2Scene: Text-driven Indoor Scene Stylization with Part-Aware DetailsInwoo Hwang, Hyeonwoo Kim, Young Min KimCVPR 2023
- VODiff: Controlling Object Visibility Order in Text-to-Image GenerationDong Liang, Jinyuan Jia, Yuhao Liu, Zhanghan Ke et al.CVPR 2025
- Segment Any 3D Object with LanguageSeungjun Lee, Yuyang Zhao, Gim Hee LeeICLR 2025
