ViewSRD: 3D Visual Grounding Via Structured Multi-View Decomposition
Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
摘要
3D visual grounding aims to identify and localize objects in a 3D space based on textual descriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsistencies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a framework that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Relation Decoupling (SRD) module restructures complex multianchor queries into a set of targeted single-anchor statements, generating a structured set of perspective-aware descriptions that clarify positional relationships. These decomposed representations serve as the foundation for the Multi-view Textual-Scene Interaction (Multi-TSI) module, which integrates textual and scene features across multiple viewpoints using shared, Cross-modal Consistent View Tokens (CCVTs) to preserve spatial correlations. Finally, a Textual-Scene Reasoning module synthesizes multi-view predictions into a unified and robust 3D visual grounding. Experiments on 3D visual grounding datasets show that ViewSRD significantly outperforms state-of-the-art methods, particularly in complex queries requiring precise spatial differentiation. Code is available at https : / / github.com/visualjason/ViewSRD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- S^2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural GuidanceBeining Xu, Siting Zhu, Zhao Jin, Junxian Li 等CVPR 2026
- ORD: Object-Relation Decoupling for Generalized 3D Visual GroundingRonggang Huang, Fansen Meng, Huaidong Zhang, Xuemiao XuCVPR 2026
它引用的顶会 Paper22
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- OpenChat: Advancing Open-source Language Models with Mixed-Quality DataGuan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li 等ICLR 2024 · 被引用 328 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- SAT: 2D Semantics Assisted Training for 3D Visual GroundingZhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo LuoICCV 2021 · 被引用 166 次
相关 Paper
- Multi-View Transformer for 3D Visual GroundingShijia Huang, Yilun Chen, Jiaya Jia, Liwei WangCVPR 2022 · 被引用 97 次
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang 等ICCV 2023 · 被引用 86 次
- MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual GroundingChun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier StrickerCVPR 2024
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu 等CVPR 2022 · 被引用 146 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
