AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
Xinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo, Xun Yang
摘要
3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly encounter a shortage: a limited amount and diversity of text-3D pairs available for training. Moreover, they fall short in effectively leveraging different contextual clues (e.g., rich spatial relations within the 3D visual space) for grounding. To address these limitations, we propose AugRefer, a novel approach for advancing 3D visual grounding. AugRefer introduces cross-modal augmentation designed to extensively generate diverse text-3D pairs by placing objects into 3D scenes and creating accurate and semantically rich descriptions using foundation models. Notably, the resulting pairs can be utilized by any existing 3DVG methods for enriching their training data. Additionally, AugRefer presents a language-spatial adaptive decoder that effectively adapts the potential referring objects based on the language description and various 3D spatial relations. Extensive experiments on three benchmark datasets clearly validate the effectiveness of AugRefer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu 等NeurIPS 2025 · 被引用 18 次
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationJie Xu, Na Zhao, Gang Niu, Masashi Sugiyama 等ICCV 2025 · 被引用 6 次
- Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language ModelsShengli Zhou, Minghang Zheng, Feng Zheng, Yang LiuCVPR 2026 · 被引用 2 次
- Few-Shot Incremental 3D Object Detection in Dynamic Indoor EnvironmentsYun Zhu, Jianjun Qian, Jian Yang, Jin Xie 等CVPR 2026 · 被引用 2 次
- Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud AnalysisShangbo Yuan, Jie Xu, Ping Hu, Xiaofeng Zhu 等AAAI 2026
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu 等ICCV 2021 · 被引用 368 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
相关 Paper
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi 等NeurIPS 2023 · 被引用 33 次
- Towards CLIP-Driven Language-Free 3D Visual Grounding via 2D-3D Relational Enhancement and ConsistencyYuqi Zhang, Han Luo, Yinjie LeiCVPR 2024 · 被引用 5 次
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual GroundingDailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui 等ACM MM 2021 · 被引用 81 次
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 被引用 157 次
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang 等ICCV 2023 · 被引用 86 次
