MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang, Fei Chen, Steven McDonagh, Gerasimos Lampouras, Ignacio Iacobacci, Sarah Parisot
摘要
Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout conditioning, or image editing techniques which often require hand drawn masks. Nonetheless, pre-existing works struggle to take advantage of the natural instance-level compositionality of scenes due to the typically flat nature of rasterized RGB output images. Towards adressing this challenge, we introduce MuLAn: a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multi-layer, instance-wise RGBA decompositions, and over 100K instance images. To build MuLAn, we developed a training free pipeline which decomposes a monocular RGB image into a stack of RGBA layers comprising of background and isolated instances. We achieve this through the use of pre-trained general-purpose models, and by developing three modules: image decomposition for instance discovery and extraction, instance completion to reconstruct occluded areas, and image re-assembly. We use our pipeline to create MuLAn-COCO and MuLAn-LAION datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance decompo-sition and occlusion information for high quality images, opening up new avenues for text-to-image generative AI re-search. With this, we aim to encourage the development of novel generation and editing technology, in particular layer-wise solutions. MuLAn data resources are available at https://MuLAn-dataset.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- PromptFix: You Prompt and We Fix the PhotoYongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu 等NeurIPS 2024 · 被引用 55 次
- Qwen-Image-Layered: Towards Inherent Editability via Layer DecompositionShengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao 等CVPR 2026 · 被引用 30 次
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu 等CVPR 2026 · 被引用 26 次
- Precise Object and Effect Removal with Adaptive Target-Aware AttentionJixin Zhao, Zhouxia Wang, Peiqing Yang, Shangchen ZhouCVPR 2026 · 被引用 13 次
- LayerFlow: A Unified Model for Layer-aware Video GenerationSihui Ji, Hao Luo, Xi Chen, Yuanpeng Tu 等SIGGRAPH 2025 · 被引用 12 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang 等NeurIPS 2024 · 被引用 13 次
- Referring Layer DecompositionFangyi Chen, Yaojie Shen, Lu Xu, Ye Yuan 等ICLR 2026 · 被引用 3 次
- ConsistCompose: Unified Multimodal Layout Control for Image CompositionXuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 等CVPR 2026 · 被引用 5 次
- Compositional Text-to-Image Generation with Dense Blob RepresentationsWeili Nie, Sifei Liu, Morteza Mardani, Chao Liu 等ICML 2024 · 被引用 44 次
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo 等CVPR 2026 · 被引用 10 次
