Lune

CVPR2026顶会

EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding

GwangWook Park, Hyo-Jun Lee, Jong-Hyeon Baek, Hanul Kim, Yeong Jun Koh

出版方
2026年份

摘要

Despite recent progress in 3D visual grounding, existing methods still struggle with three core challenges: 1) crossmodal misalignment that prevents textual cues from being reliably delivered to visual representations, 2) intra-class confusion arising from insufficient understanding of finegrained expression cues, and 3) geometric reasoning errors caused by inaccurate aggregation of spatially relevant visual features. We propose EG-3DVG, a unified framework that addresses these issues through an expression and geometry aware grounding decoder. The decoder integrates two complementary attention modules-position-guided expression cross-attention (PECA) for reliable text-vision alignment and geometry-aware masked attention (GMA) for selective aggregation of geometry-consistent visual cues. To further distinguish semantically similar instances, we introduce expression-aware contrastive learning (ECL), which strengthens the alignment between the target object token and expression-relevant words. Extensive experiments on ScanRefer and SR3D/NR3D demonstrate that EG-3DVG achieves state-of-the-art performance in both 3D bounding box localization and mask prediction, validating the effectiveness of our geometry-and expression-aware design. Code is available at https://github.com/ Gwan9Wook/EG3DVG.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 45d82f39-72f7-4749-8102-e4c477232ffc

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖