Muses: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
Yanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue, Yu Qiao, Yali Wang
Abstract
Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system, namely MUSES, for 3D-controllable image generation from user queries. Specifically, our MUSES develops a progressive workflow with three key components, including (1) Layout Manager for 2D-to-3D layout lifting, (2) Model Engineer for 3D object acquisition and calibration, (3) Image Artist for 3D-to-2D image rendering. By mimicking the collaboration of human professionals, this multi-modal agent pipeline facilitates the effective and automatic creation of images with 3D-controllable objects, through an explainable integration of top-down planning and bottom-up generation. Additionally, existing benchmarks lack detailed descriptions of complex 3D spatial relationships of multiple objects. To fill this gap, we further construct a new benchmark of T2I-3DisBench (3D image scene), which describes diverse 3D image scenes with 50 detailed prompts. Extensive experiments show the state-of-the-art performance of MUSES on both T2I-CompBench and T2I-3DisBench, outperforming recent strong competitors such as DALL-E 3 and Stable Diffusion 3. These results demonstrate a significant step forward for MUSES in bridging natural language, 2D image generation, and 3D world.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- MTVCraft: Tokenizing 4D Motion for Arbitrary Character AnimationYanbo Ding, Xirui Hu, Guo Zhi, Yan Zhang et al.ICLR 2026 · 3 citations
- DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial EditabilityXirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang et al.ICCV 2025 · 3 citations
- MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image AnimationXirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi et al.ICLR 2026 · 3 citations
- 3D Space as a Scratchpad for Editable Text-to-Image GenerationOindrila Saha, Vojtech Krs, Radomír Mech, Subhransu Maji et al.CVPR 2026 · 1 citation
- V-Stylist: Video Stylization via Collaboration and Reflection of MLLM AgentsZhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding et al.CVPR 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic ExpansionFei Peng, Junqiang Wu, Yan Li, Tingting Gao et al.ICCV 2025 · 1 citation
- Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image GenerationAbdelrahman Eldesokey, Peter WonkaICLR 2025 · 1 citation
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang et al.NeurIPS 2024 · 13 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- MIDI: Multi-Instance Diffusion for Single Image to 3D Scene GenerationZehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang et al.CVPR 2025
