3D Question Answering via only 2D Vision-Language Models
Fengyun Wang, Sicheng Yu, Jiawei Wu, Jinhui Tang, Hanwang Zhang, Qianru Sun
摘要
Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, using 3D question answering (3D-QA) as a representative example. Due to the limited training data in 3D, we do not train LVLMs but infer in a zero-shot manner. Specifically, we sample 2D views from a 3D point cloud and feed them into 2D models to answer a given question. When the 2D model is chosen, e.g., LLAVA-OV, the quality of sampled views matters the most. We propose cdViews, a novel approach to automatically selecting critical and diverse Views for 3D-QA. cdViews consists of two key components: viewSelector prioritizing critical views based on their potential to provide answer-specific information, and viewNMS enhancing diversity by removing redundant views based on spatial overlap. We evaluate cdViews on the widely-used ScanQA and SQA benchmarks, demonstrating that it achieves state-of-the-art performance in 3D-QA while relying solely on 2D models without fine-tuning. These findings support our belief that 2D LVLMs are currently the most effective alternative (of the resource-intensive 3D LVLMs) for addressing 3D tasks. The code is available at https: //github.com/fereenwong/cdViews .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token MergingTianbo Pan, Xingyi Yang, Xinchao WangCVPR 2026
- Scalable Multi-Task Low-Rank Model AdaptationZichen Tian, Antoine Ledent, Qianru SunICLR 2026
- Zero-Shot 3D Question Answering via Hierarchical View-to-Token TransportationDongsheng Wang, Dawei Su, Hui HuangICML 2026
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke 等CVPR 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu 等ICML 2024 · 被引用 361 次
相关 Paper
- LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMsDoriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot 等AAAI 2026
- Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQAWentao Mo, Yang LiuAAAI 2024 · 被引用 30 次
- Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene UnderstandingDuo Zheng, Shijia Huang, Liwei WangCVPR 2025
- LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningSijin Chen, Xin Chen, Chi Zhang, Mingsheng Li 等CVPR 2024
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas 等ICCV 2025 · 被引用 4 次
