CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians
Chongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng, Masayoshi Tomizuka, Ping Luo, Mingyu Ding, Varun Jampani, Wei Zhan
Abstract
Adobe Research 4 UNC-Chapel Hill 5 Stability AI an owl an owl perches on a branch an owl perches on a branch near a pinecone an owl perches on a branch near a pinecone, with a rat below the branch Single 3D Generation Compositional 3D Generation Rendered Images Gaussians Initialization with 2D Compositionality 1K iterations 4K iterations "an owl perches on a branch near a pinecone" a footballer is kicking a soccer ball a bird is drinking water from a cup a photographer is capturing a beautiful butterfly with camera a beautiful butterfly … Figure 1. Illustration of compositional 3D Generation and COMPGS. All the contents are generated by COMPGS. Top row: COMPGS is capable of generating either a single object (e.g., a butterfly) or generating compositional objects with reasonable interactions (e.g., the rightmost figure in the top row). Middle row: Beyond text-to-3D generation, COMPGS can be easily extend to 3D editing by progressively adding objects. The colored texts (e.g., 'a branch', 'a pinecone', 'a rat' in the rightmost figure) denote the added part compared to its previous asset. Bottom row: COMPGS achieves compositional text-to-3D by transferring 2D compositionality to initialize 3D Gaussians. COMPGS is further trained with dynamic SDS optimization to produce plausible results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content CreationSankalp Sinha, Mohammad Sadil Khan, Muhammad Usama, Shino Sam et al.CVPR 2025
- CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D CreationBonan Li, Zicheng Zhang, Xingyi Yang, Xinchao WangCVPR 2025
- A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D SupervisionChensheng Peng, Ido Sobol, Masayoshi Tomizuka, Kurt Keutzer et al.ICCV 2025
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Multitwine: Multi-Object Compositing with Text and Layout ControlGemma Canet Tarrés, Zhe Lin, Zhifei Zhang, He Zhang et al.CVPR 2025
- Muses: 3D-Controllable Image Generation via Multi-Modal Agent CollaborationYanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue et al.AAAI 2025 · 8 citations
- TextCraftor: Your Text Encoder can be Image Quality ControllerYanyu Li, Xian Liu, Anil Kag, Ju Hu et al.CVPR 2024
- Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D GenerationYiming Qin, Zhu Xu, Yang LiuCVPR 2025
- Layout-your-3D: Controllable and Precise 3D Generation with 2D BlueprintJunwei Zhou, Xueting Li, Lu Qi, Ming-Hsuan YangICLR 2025
