Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
Wentao Mo, Yang Liu
摘要
In 3D Visual Question Answering (3D VQA), the scarcity of fully annotated data and limited visual content diversity hampers the generalization to novel scenes and 3D concepts (e.g., only around 800 scenes are utilized in ScanQA and SQA dataset). Current approaches resort supplement 3D reasoning with 2D information. However, these methods face challenges: either they use top-down 2D views that introduce overly complex and sometimes question-irrelevant visual clues, or they rely on globally aggregated scene/imagelevel representations from 2D VLMs, losing the fine-grained vision-language correlations. To overcome these limitations, our approach utilizes question-conditional 2D view selection procedure, pinpointing semantically relevant 2D inputs for crucial visual clues. We then integrate this 2D knowledge into the 3D-VQA system via a two-branch Transformer structure. This structure, featuring a Twin-Transformer design, compactly combines 2D and 3D modalities and captures fine-grained correlations between modalities, allowing them mutually augmenting each other. Integrating proposed mechanisms above, we present BridgeQA, that offers a fresh perspective on multi-modal transformer-based architectures for 3D-VQA. Experiments validate that BridgeQA achieves state-of-the-art on 3D-VQA datasets and significantly outperforms existing solutions. Code is available at https://github . com/matthewdm0816/BridgeQA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Qi Chen, Jiayi Ji, Haowei Wang 等NeurIPS 2024 · 被引用 16 次
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng 等ICCV 2025 · 被引用 6 次
- ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingMinghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng 等ACM MM 2024 · 被引用 5 次
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge TransferringZhu Xu, Ting Lei, Zhimin Li, Guan Wang 等ICCV 2025 · 被引用 3 次
- Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language ModelsShengli Zhou, Minghang Zheng, Feng Zheng, Yang LiuCVPR 2026 · 被引用 2 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai 等ICLR 2022 · 被引用 950 次
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 被引用 135 次
相关 Paper
- 3D Question Answering with Scene Graph ReasoningZizhao Wu, Haohan Li, Gongyi Chen, Zhou Yu 等ACM MM 2024 · 被引用 6 次
- 3D Question Answering via only 2D Vision-Language ModelsFengyun Wang, Sicheng Yu, Jiawei Wu, Jinhui Tang 等ICML 2025
- DSPNet: Dual-vision Scene Perception for Robust 3D Question AnsweringJingzhou Luo, Yang Liu, Weixing Chen, Zhen Li 等CVPR 2025
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 被引用 48 次
