SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
Xiongkun Linghu, Jiangyong Huang, Ziyu Zhu, Baoxiong Jia, Siyuan Huang
摘要
Existing research on 3D Large Language Models (LLMs) still struggles to achieve grounded question-answering, primarily due to the under-exploration of the mechanism of human-like scene-object grounded reasoning. This paper bridges the gap by presenting a novel framework. We first introduce a grounded Chain-of-Thought reasoning method in 3D scenes (SCENECOT), decoupling a complex reasoning task into simpler and manageable problems, and building corresponding visual clues based on multimodal expert modules. To enable such a method, we develop SCENECOT-185K, the first large-scale grounded CoT reasoning dataset, consisting of 185K high-quality instances. Extensive experiments across various complex 3D scene reasoning benchmarks demonstrate that our new framework achieves strong performance with high grounding-QA coherence. To the best of our knowledge, this is the first successful application of CoT reasoning to 3D scene understanding, enabling step-by-step human-like reasoning and showing potential for extension to broader 3D scene understanding scenarios. Code and data are available on project page. INTRODUCTION Understanding 3D scenes is a fundamental capability for building human-level embodied agents (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token MergingTianbo Pan, Xingyi Yang, Xinchao WangCVPR 2026
- PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-ThoughtChaoqi Chen, Qile Xu, Wenjun Zhou, Hui HuangSIGGRAPH 2026
它引用的顶会 Paper25
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
相关 Paper
- SCoT: Teaching 3D-LLMs to Think Spatially with Million-scale CoT AnnotationsJinpeng Li, Haiping Wang, Jiabin Chen, Yuan Liu 等ICLR 2026
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu 等NeurIPS 2025 · 被引用 18 次
- GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language InstructionsXiaomeng Chu, Jiajun Deng, Guoliang You, Wei Liu 等ICCV 2025
- CoT-VLNBench: A Benchmark for Visual Chain-of-Thought Reasoning in Vision-Language-Navigation RobotsXiao Zhao, Chang Liu, Ruiteng Ji, Zheyuan Zhang 等AAAI 2026
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang 等ICML 2024 · 被引用 182 次
