G3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual Grounding
Yuan Wang, Yali Li, Shengjin Wang
摘要
Grounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI, as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counterparts, challenges posed by the variability of 3D visual grounding remain relatively unsolved in existing studies: 1) the underlying geometric and complex spatial relationships in 3D scene. 2) the inherent complexity of 3D grounded language. 3) the inconsistencies between text and geometric features. To tackle these issues, we propose G<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>-LQ, a DEtection TRansformer-based model tailored for 3D visual grounding task. G<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>-LQ explicitly models Geometric-aware visual rep-resentations and Generates fine-Grained Language-guided object Queries in an overarching framework, which com-prises two dedicated modules. Specifically, the Position Adaptive Geometric Exploring (PAGE) unearths underlying information of 3D objects in the geometric details and spatial relationships perspectives. The Fine-grained Language-guided Query Selection (Flan-QS) delves into syntactic structure of texts and generates object queries that exhibit higher relevance towards fine-grained text features. Finally, a pioneering Poincaré Semantic Alignment (PSA) loss establishes semantic-geometry consistencies by modeling non-linear vision-text feature mappings and aligning them on a hyperbolic prototype-Poincaré ball. Extensive experiments verify the superiority of our G<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>-LQ method, trumping the state-of-the-arts by a considerable margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual GroundingYuan Wang, Yali Li, Eastman Z. Y. Wu, Shengjin WangAAAI 2025 · 被引用 11 次
- SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual GroundingJiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan 等ACM MM 2025 · 被引用 4 次
- Modality Alignment across Trees on Heterogeneous Hyperbolic ManifoldsWei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao 等ICLR 2026 · 被引用 3 次
- ViewSRD: 3D Visual Grounding Via Structured Multi-View DecompositionRonggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu 等ICCV 2025 · 被引用 2 次
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingYihang Zhu, Jinhao Zhang, Yuxuan Wang, Aming Wu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper30
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual GroundingDailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui 等ACM MM 2021 · 被引用 81 次
- Multi-View Transformer for 3D Visual GroundingShijia Huang, Yilun Chen, Jiaya Jia, Liwei WangCVPR 2022 · 被引用 97 次
- MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual GroundingChun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier StrickerCVPR 2024
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu 等CVPR 2022 · 被引用 146 次
