Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
Wei Deng, Mengshi Qi, Huadong Ma
摘要
Large Vision-Language Models (VLMs), such as GPT-4, have achieved remarkable success across various fields. However, there are few studies on 3D indoor scene generation with VLMs. This paper considers this task as a planning problem subject to spatial and layout common sense constraints. To solve the problem with a VLM, we propose a new global-local tree search algorithm. Globally, the method places each object sequentially and explores multiple placements during each placement process, where the problem space is represented as a tree. To reduce the depth of the tree, we decompose the scene structure hierarchically, i.e. room level, region level, floor object level, and supported object level. The algorithm independently generates the floor objects in different regions and supported objects placed on different floor objects. Locally, we also decompose the sub-task, the placement of each object, into multiple steps. The algorithm searches the tree of problem space. To leverage the VLM model to produce positions of objects, we discretize the top-down view space as a dense grid and fill each cell with diverse emojis to make to cells distinct. We prompt the VLM with the emoji grid and the VLM produces a reasonable location for the object by describing the position with the name of emojis. The quantitative and qualitative experimental results illustrate our approach generates more plausible 3D scenes than stateof-the-art approaches. Our source code is available at https://github.com/dw-dengwei/TreeSearchGen .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial ReasoningXingjian Ran, Yixuan Li, Linning Xu, Mulin Yu 等NeurIPS 2025 · 被引用 34 次
- FlashWorld: High-quality 3D Scene Generation within SecondsXinyang Li, Tengfei Wang, Zixiao Gu, Shengchuan Zhang 等ICLR 2026 · 被引用 32 次
- Video Perception Models for 3D Scene SynthesisRui Huang, Guangyao Zhai, Zuria Bauer, Marc Pollefeys 等NeurIPS 2025 · 被引用 12 次
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 被引用 12 次
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene ComplexityHaoming Wang, Qiyao Xue, Wei GaoCVPR 2026 · 被引用 6 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang 等ICLR 2026 · 被引用 121 次
- Hierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language ModelWeilin Sun, Xinran Li, Manyi Li, Kai Xu 等AAAI 2025 · 被引用 7 次
- FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object PlacementIan Huang, Yanan Bao, Karen Truong, Howard Zhou 等CVPR 2025
- Holodeck: Language Guided Generation of 3D Embodied AI EnvironmentsYue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt 等CVPR 2024 · 被引用 47 次
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 被引用 423 次
