Towards CLIP-Driven Language-Free 3D Visual Grounding via 2D-3D Relational Enhancement and Consistency
Yuqi Zhang, Han Luo, Yinjie Lei
Abstract
3D visual grounding plays a crucial role in scene understanding, with extensive applications in AR/VR. Despite the significant progress made in recent methods, the re-quirement of dense textual descriptions for each individ-ual object, which is time-consuming and costly, hinders their scalability. To mitigate reliance on text annotations during training, researchers have explored language-free training paradigms in the 2D field via explicit text gen-eration or implicit feature substitution. Nevertheless, un-like 2D images, the complexity of spatial relations in 3D, coupled with the absence of robust 3D visual language pre-trained models, makes it challenging to directly trans-fer previous strategies. To tackle the above issues, in this paper, we introduce a language-free training framework for 3D visual grounding. By utilizing the visual-language joint embedding in 2D large cross-modality model as a bridge, we can expediently produce the pseudo-language features by leveraging the features of 2D images which are equivalent to that of real textual descriptions. We fur-ther develop a relation injection scheme, with a Neighboring Relation-aware Modeling module and a Cross-modality Relation Consistency module, aiming to enhance and pre-serve the complex relationships between the 2D and 3D embedding space. Extensive experiments demonstrate that our proposed language-free 3D visual grounding approach can obtain promising performance across three widely used datasets - ScanRefer, Nr3D and Sr3D. Our codes are avail-able at https://github.com/xibi777/3DLFVG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dc7129b-3f1b-4ac3-8a10-ddc365be9013Cited by top-tier papers10
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual GroundingZhao Jin, Rong-Cheng Tu, Jingyi Liao, Wenhao Sun et al.NeurIPS 2025 · 13 citations
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng et al.NeurIPS 2025 · 6 citations
- SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual GroundingJiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan et al.ACM MM 2025 · 4 citations
- ViewSRD: 3D Visual Grounding Via Structured Multi-View DecompositionRonggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu et al.ICCV 2025 · 2 citations
- Perspective from a Broader Context: Can Room Style Knowledge Help Visual Floorplan Localization?Bolei Chen, Shengsheng Yan, Yongzheng Cui, Jiaxu Kang et al.AAAI 2026 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo et al.AAAI 2025 · 12 citations
- Multi-Attribute Interactions Matter for 3D Visual GroundingCan Xu, Yuehui Han, Rui Xu, Le Hui et al.CVPR 2024 · 5 citations
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang et al.ICCV 2023 · 86 citations
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang et al.CVPR 2026
