NS3D: Neuro-Symbolic Grounding of 3D Objects and Relations
Joy Hsu, Jiayuan Mao, Jiajun Wu
摘要
Grounding object properties and relations in 3D scenes is a prerequisite for a wide range of artificial intelligence tasks, such as visually grounded dialogues and embodied manipulation. However, the variability of the 3D domain induces two fundamental challenges: 1) the expense of labeling and 2) the complexity of 3D grounded language. Hence, essential desiderata for models are to be data-efficient, generalize to different data distributions and tasks with unseen semantic forms, as well as ground complex language semantics (e.g., view-point anchoring and multi-object reference). To address these challenges, we propose NS3D, a neuro-symbolic framework for 3D grounding. NS3D translates language into programs with hierarchical structures by leveraging large language-to-code models. Different functional modules in the programs are implemented as neural networks. Notably, NS3D extends prior neuro-symbolic visual reasoning methods by introducing functional modules that effectively reason about high-arity relations (i.e., relations among more than two objects), key in disambiguating objects in complex 3D scenes. Modular and compositional architecture enables NS3D to achieve state-of-the-art results on the ReferIt3D view-dependence task, a 3D referring expression comprehension benchmark. Importantly, NS3D shows significantly improved performance on settings of data-efficiency and generalization, and demonstrate zero-shot transfer to an unseen 3D question-answering task. * For conciseness, we have used filter(shelf) as a short-hand notation for filter(scene(), shelf).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 被引用 157 次
- What's Left? Concept Grounding with Logic-Enhanced Foundation ModelsJoy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, Jiajun WuNeurIPS 2023 · 被引用 54 次
- Motion Question Answering via Modular Motion ProgramsMark Endo, Joy Hsu, Jiaman Li, Jiajun WuICML 2023 · 被引用 28 次
- Visual Programming for Zero-Shot Open-Vocabulary 3D Visual GroundingZhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao 等CVPR 2024 · 被引用 19 次
- CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual GroundingEslam Mohamed Bakr, Mohamed Ayman, Mahmoud Ahmed, Habib Slim 等ICLR 2024 · 被引用 16 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak 等ICCV 2019 · 被引用 474 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- Synchromesh: Reliable Code Generation from Pre-trained Language ModelsGabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari 等ICLR 2022 · 被引用 200 次
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 被引用 191 次
相关 Paper
- Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision GroundingJiaxin Shi, Mingyue Xiang, Hao Sun, Yixuan Huang 等CVPR 2025
- Naturally Supervised 3D Visual Grounding with Language-Regularized Concept LearnersChun Feng, Joy Hsu, Weiyu Liu, Jiajun WuCVPR 2024
- Towards CLIP-Driven Language-Free 3D Visual Grounding via 2D-3D Relational Enhancement and ConsistencyYuqi Zhang, Han Luo, Yinjie LeiCVPR 2024 · 被引用 5 次
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo 等AAAI 2025 · 被引用 12 次
- MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual GroundingChun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier StrickerCVPR 2024
