SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang
摘要
The window to the right of the oven hood." "The laptop beside the floral-patterned chair." "The bookshelf second from the right. " "The chair backed to the window." "The closed door. Not the bathroom door. " "Select the couch that has an L shape." Previous SoTA Ours (a) Texture (b) Shape (c) Viewpoint (d) Orientation (e) State (f) Order Closed Open Right Left Right Illustration Flor. chair Chair back L shape I shape Window Window Lap. 1 st 3 rd 4 th 2 nd Window Oven hood Figure 1. Effectiveness of SeeGround: Different from previous SoTA, our method associates 2D visual cues -color, texture, viewpoint, spatial position, orientation, and state -with 3D spatial text description to achieve precise scene understanding. Specifically, our method: (a) identifies the floral chair by recognizing unique color and texture cues; (b) recognizes the couch by interpreting geometric shape; (c) determines the right window by interpreting spatial relationships and perspective; (d) identifies the chair by discerning directional alignment; (e) detects the closed door by visually interpreting its state; and (f) selects the bookshelf by understanding relative positioning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong 等ICLR 2026 · 被引用 28 次
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual GroundingZhao Jin, Rong-Cheng Tu, Jingyi Liao, Wenhao Sun 等NeurIPS 2025 · 被引用 13 次
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR SequencesXiang Xu, Ao Liang, Youquan Liu, Linfeng Li 等CVPR 2026 · 被引用 8 次
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li 等NeurIPS 2025 · 被引用 7 次
- 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive DecodingMakanjuola Adekunmi Ogunleye, Eman Abdelrahman, Ismini LourentzouCVPR 2026 · 被引用 3 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 等ICCV 2023 · 被引用 620 次
相关 Paper
- OpenScene: 3D Scene Understanding with Open VocabulariesSongyou Peng, Kyle Genova, Chiyu Max Jiang, Andrea Tagliasacchi 等CVPR 2023
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingHongjia Zhai, Hai Li, Zhenzhe Li, Xiaokun Pan 等CVPR 2025
- Text2Scene: Text-driven Indoor Scene Stylization with Part-Aware DetailsInwoo Hwang, Hyeonwoo Kim, Young Min KimCVPR 2023
- VODiff: Controlling Object Visibility Order in Text-to-Image GenerationDong Liang, Jinyuan Jia, Yuhao Liu, Zhanghan Ke 等CVPR 2025
- Segment Any 3D Object with LanguageSeungjun Lee, Yuyang Zhao, Gim Hee LeeICLR 2025
