3D Concept Learning and Reasoning from Multi-View Images
Yining Hong, Chunru Lin, Yilun Du, Zhenfang Chen, Joshua B. Tenenbaum, Chuang Gan
摘要
Humans are able to accurately reason in 3D by gathering multi-view observations of the surrounding world. Inspired by this insight, we introduce a new large-scale benchmark for 3D multi-view visual question answering (3DMV-VQA). This dataset is collected by an embodied agent actively moving and capturing RGB images in an environment using the Habitat simulator. In total, it consists of approximately 5k scenes, 600k images, paired with 50k questions. We evaluate various state-of-the-art models for visual reasoning on our benchmark and find that they all perform poorly. We suggest that a principled approach for 3D reasoning from multi-view images should be to infer a compact 3D representation of the world from the multi-view images, which is further grounded on open-vocabulary semantic concepts, and then to execute reasoning on these 3D representations. As the first step towards this approach, we propose a novel 3D concept learning and reasoning (3D-CLR) framework that seamlessly combines these components via neural fields, 2D pre-trained vision-language models, and neural reasoning operators. Experimental results suggest that our framework outperforms baseline models by a large margin, but the challenge remains largely unsolved. We further perform an in-depth analysis of the challenges and highlight potential future directions. .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- Cameras as Relative Positional EncodingRuilong Li, Brent Yi, Junchen Liu, Hang Gao 等NeurIPS 2025 · 被引用 113 次
- Spatial Mental Modeling from Limited ViewsQineng Wang, Baiqiao Yin, Pingyue Zhang, Jianshu Zhang 等ICLR 2026 · 被引用 92 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li 等ICCV 2021 · 被引用 1,284 次
相关 Paper
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsWeijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao 等NeurIPS 2025 · 被引用 4 次
- SQA3D: Situated Question Answering in 3D ScenesXiaojian Ma, Silong Yong, Zilong Zheng, Qing Li 等ICLR 2023 · 被引用 16 次
- Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 等CVPR 2026
- Spatial Understanding from Videos: Structured Prompts Meet Simulation DataHaoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen 等NeurIPS 2025 · 被引用 31 次
- Hypo3D: Exploring Hypothetical Reasoning in 3DYe Mao, Weixun Luo, Junpeng Jing, Anlan Qiu 等ICML 2025
