Viewpoint-Aware Visual Grounding in 3D Scenes
Xiangxi Shi, Zhonghua Wu, Stefan Lee
Abstract
Referring expressions for visual objects often include descriptions of relative spatial arrangements to other objects - e.g. “to the right of” - that depend on the point of view of the speaker. In 2D referring expression tasks, this view-point is captured unambiguously in the image. However, grounding expressions with such spatial language in 3D without viewpoint annotations can be ambiguous. In this paper, we investigate the significance of viewpoint information in 3D visual grounding - introducing a model that explicitly predicts the speaker's viewpoint based on the referring expression and scene. We pretrain this model on a synthetically generated dataset that provides viewpoint annotations and then finetune on 3D referring expression datasets. Further, we introduce an auxiliary uniform object representation loss to encourage viewpoint invariance in learned object representations. We find that our proposed ViewPoint Prediction Network (VPP-Net) achieves state-of-the-art performance on ScanRefer, SR3D, and NR3D - improving Accuracy@0.25IoU by 1.06%, 0.60%, and 2.00% respectively compared to prior work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7504bfa2-d7e6-4f22-8cd4-1bcd2752d87bCited by top-tier papers9
- Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language ModelsShengli Zhou, Minghang Zheng, Feng Zheng, Yang LiuCVPR 2026 · 2 citations
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingYihang Zhu, Jinhao Zhang, Yuxuan Wang, Aming Wu et al.ICCV 2025 · 1 citation
- Text-guided Sparse Voxel Pruning for Efficient 3D Visual GroundingWenxuan Guo, Xiuwei Xu, Ziwei Wang, Jianjiang Feng et al.CVPR 2025
- Promptable 3-D Object Localization with Latent Diffusion ModelsCheng-Yao Hong, Li-Heng Wang, Tyng-Luh LiuNeurIPS 2025
- EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual GroundingGwangWook Park, Hyo-Jun Lee, Jong-Hyeon Baek, Hanul Kim et al.CVPR 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
Related papers
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.EMNLP 2023 · 7 citations
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
- Refer-It-in-RGBD: A Bottom-Up Approach for 3D Visual Grounding in RGBD ImagesHaolin Liu, Anran Lin, Xiaoguang Han, Lei Yang et al.CVPR 2021
- Multi-View Transformer for 3D Visual GroundingShijia Huang, Yilun Chen, Jiaya Jia, Liwei WangCVPR 2022 · 97 citations
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang et al.ICCV 2023 · 86 citations
