Transparent Image Layer Diffusion using Latent Transparency
Lvmin Zhang, Maneesh Agrawala
摘要
We present an approach enabling large-scale pretrained latent diffusion models to generate transparent images. The method allows generation of single transparent images or of multiple transparent layers. The method learns a "latent transparency" that encodes alpha channel transparency into the latent manifold of a pretrained latent diffusion model. It preserves the production-ready quality of the large diffusion model by regulating the added transparency as a latent offset with minimal changes to the original latent distribution of the pretrained model. In this way, any latent diffusion model can be converted into a transparent image generator by finetuning it with the adjusted latent space. We train the model with 1M transparent image layer pairs collected using a human-in-the-loop collection scheme. We show that latent transparency can be applied to different open source image generators, or be adapted to various conditional control systems to achieve applications like foreground/background-conditioned layer generation, joint layer generation, structural control of layer contents, etc. A user study finds that in most cases (97%) users prefer our natively generated transparent content over previous ad-hoc solutions such as generating and then matting. Users also report the quality of our generated transparent images is comparable to real commercial transparent assets like Adobe Stock.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper58
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic DesignHui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng 等ICLR 2026 · 被引用 40 次
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li 等NeurIPS 2025 · 被引用 30 次
- Qwen-Image-Layered: Towards Inherent Editability via Layer DecompositionShengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao 等CVPR 2026 · 被引用 30 次
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu 等CVPR 2026 · 被引用 26 次
- LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene GenerationShuai Yang, Jing Tan, Mengchen Zhang, Tong Wu 等SIGGRAPH 2025 · 被引用 23 次
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Matting by GenerationZhixiang Wang, Baiang Li, Jian Wang, Yu-Lun Liu 等SIGGRAPH 2024 · 被引用 5 次
- DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image MattingXiaodi Li, Zongxin Yang, Ruijie Quan, Yi YangNeurIPS 2024 · 被引用 15 次
- Trans-Adapter: A Plug-And-Play Framework for Transparent Image InpaintingYuekun Dai, Haitian Li, Shangchen Zhou, Chen Change LoyICCV 2025 · 被引用 1 次
- Video Generation with Stable Transparency via Shiftable RGB-A Distribution LearnerHaotian Dong, Wenjing Wang, Chen Li, Jing LYU 等CVPR 2026 · 被引用 7 次
- From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer DecompositionJingxi Chen, Yixiao Zhang, Xiaoye Qian, Zongxia Li 等CVPR 2026 · 被引用 5 次
