Splattalk: 3D VQA with Gaussian Splatting
Anh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas, Thomas A. Funkhouser
摘要
Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D VQA tasks, progress in the 3D domain has been significantly slower due to the complexity of 3D data and the high cost of manual annotations. In this work, we introduce SplatTalk, a novel method that uses a generalizable 3D Gaussian Splatting (3DGS) framework to produce 3D tokens suitable for direct input into a pretrained LLM, enabling effective zero-shot 3D visual question answering (3D VQA) for scenes with only posed images. During experiments on multiple benchmarks, our approach outperforms both 3D models trained specifically for the task and previous 2D-LMM-based models utilizing only images (our setting), while achieving competitive performance with state-of-the-art 3D LMMs that additionally utilize 3D inputs. Project website: https://splat-talk.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial ReasoningJian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi 等CVPR 2026 · 被引用 16 次
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D ScenesXiongkun Linghu, Jiangyong Huang, Ziyu Zhu, Baoxiong Jia 等ICLR 2026 · 被引用 9 次
- ReLaGS: Relational Language Gaussian SplattingYaxu Xie, Abdalla Arafa, Alireza Javanmardi, Christen Millerdurai 等CVPR 2026 · 被引用 7 次
- Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and AlignmentJerry Jiang, Haowen Sun, Denis A. Gudovskiy, Yohei Nakata 等CVPR 2026 · 被引用 3 次
- Point Cloud as a Foreign Language for Multi-modal Large Language ModelSneha Paul, Zachary Patterson, Nizar BouguilaCVPR 2026 · 被引用 2 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke 等CVPR 2026
- MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and EditingJingqiao Xiu, Can Wang, Dong XuCVPR 2026
- Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language UnderstandingYutao Tang, Cheng Zhao, Gaurav Mittal, Rohith Kukkala 等CVPR 2026 · 被引用 1 次
- LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMsDoriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot 等AAAI 2026
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
