Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
Chun Feng, Joy Hsu, Weiyu Liu, Jiajun Wu
摘要
3D visual grounding is a challenging task that often requires direct and dense supervision, notably the semantic label for each object in the scene. In this paper, we instead study the naturally supervised setting that learns from only 3D scene and QA pairs, where prior works underperform. We propose the Language-Regularized Concept Learner (LARC), which uses constraints from language as regularization to significantly improve the accuracy of neurosymbolic concept learners in the naturally supervised setting. Our approach is based on two core insights: the first is that language constraints (e.g., a word's relation to another) can serve as effective regularization for structured representations in neuro-symbolic models; the second is that we can query large language models to distill such constraints from language properties. We show that LARC improves performance of prior works in naturally supervised 3D visual grounding, and demonstrates a wide range of 3D visual reasoning capabilities-from zero-shot composition, to data efficiency and transferability. Our method represents a promising step towards regularizing structured visual reasoning frameworks with language-based priors, for learning in settings without dense supervision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential GroundingZijun Lin, Shuting He, Cheston Tan, Bihan WenICCV 2025 · 被引用 1 次
- Language-to-Space Programming for Training-Free 3D Visual GroundingBoyu Mi, Hanqing Wang, Tai Wang, Yilun Chen 等EMNLP 2025
- Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision GroundingJiaxin Shi, Mingyue Xiang, Hao Sun, Yixuan Huang 等CVPR 2025
- DenseGrounding: Improving Dense Language-Vision Semantics for Ego-centric 3D Visual GroundingHenry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng 等ICLR 2025
- Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMsWentao Mo, Yang LiuICML 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 被引用 191 次
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang 等ICCV 2021 · 被引用 188 次
相关 Paper
- NS3D: Neuro-Symbolic Grounding of 3D Objects and RelationsJoy Hsu, Jiayuan Mao, Jiajun WuCVPR 2023
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang 等CVPR 2026
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual GroundingZhao Jin, Rong-Cheng Tu, Jingyi Liao, Wenhao Sun 等NeurIPS 2025 · 被引用 13 次
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo 等AAAI 2025 · 被引用 12 次
- 3D Concept Grounding on Neural FieldsYining Hong, Yilun Du, Chunru Lin, Josh Tenenbaum 等NeurIPS 2022 · 被引用 24 次
