VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
Zhengfei Kuang, Rui Lin, Long Zhao, Gordon Wetzstein, Saining Xie, Sanghyun Woo
Abstract
Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. In this paper, we bridge this critical gap by tackling three key challenges in 3D object arrangement task using MLLMs. First, to address the weak visual grounding of MLLMs, which struggle to link programmatic edits with precise 3D outcomes, we introduce an MCP-based API. This shifts the interaction from brittle raw code manipulation to more robust, function-level updates. Second, we augment the MLLM's 3D scene understanding with a suite of specialized visual tools to analyze scene state, gather spatial information, and validate action outcomes. This perceptual feedback loop is critical for closing the gap between language-based updates and precise 3D-aware manipulation. Third, to manage the iterative, error-prone updates, we propose a collaborative multi-agent framework with designated roles for planning, execution, and verification. This decomposition allows the system to robustly handle multi-step instructions and recover from intermediate errors. We demonstrate the effectiveness of our approach on a diverse set of 25 complex object arrangement tasks, where it significantly outperforms existing baselines. Website: vulcan-3d.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 469e75de-989f-4630-8e11-8d2862a5399eBuilds on35
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi et al.ICLR 2024 · 813 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- Instruct-NeRF2NeRF: Editing 3D Scenes with InstructionsAyaan Haque, Matthew Tancik, Alexei A. Efros, Aleksander Holynski et al.ICCV 2023 · 544 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
Related papers
- FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object PlacementIan Huang, Yanan Bao, Karen Truong, Howard Zhou et al.CVPR 2025
- Visual Programming for Zero-Shot Open-Vocabulary 3D Visual GroundingZhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao et al.CVPR 2024 · 19 citations
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- Placeit3d: Language-Guided Object Placement in Real 3D ScenesAhmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi et al.ICCV 2025 · 11 citations
- GENMAC: Compositional Text-to-Video Generation with Multi-Agent CollaborationKaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin et al.AAAI 2026 · 1 citation
