DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
Zitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu, Xiangtai Li, Hao Yang, Bo Yang, Yiren Song
摘要
Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask prior dependencies, static object assumptions, and the lack of datasets. In this paper, we delve into a novel task: Layer-Wise Decomposition of Alpha-Composited Images, aiming to recover constituent layers from single overlapped images under the condition of semi-transparent/transparent alpha layer non-linear occlusion. To address challenges in layer ambiguity, generalization, and data scarcity, we first introduce AlphaBlend, the first large‑scale and high-quality dataset for transparent and semi‑transparent layer decomposition, containing six subtasks with different characteristics (e.g., translucent flare removal, semi-transparent cell decomposition, glassware decomposition). Building on this dataset, we present DiffDecompose, a diffusion Transformer-based framework that learns the posterior over possible layer decompositions conditioned on the input image, semantic prompts, and blending type. Rather than regressing alpha mattes directly, DiffDecompose performs In‑Context Decomposition, enabling the model to predict one or multiple layers without per‑layer supervision, and introduces Layer Position Encoding Cloning to maintain pixel‑level correspondence across layers. Extensive experiments on the proposed AlphaBlend dataset and public LOGO dataset verify the effectiveness of DiffDecompose. The code and dataset will be available upon paper acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li 等NeurIPS 2025 · 被引用 30 次
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive AlignmentZiheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu 等CVPR 2026 · 被引用 6 次
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen 等ICCV 2025 · 被引用 5 次
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerYiren Song, Danze Chen, Mike Zheng ShouICCV 2025 · 被引用 5 次
- FonTS: Text Rendering with Typography and Style ControlsWenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer DecompositionJingxi Chen, Yixiao Zhang, Xiaoye Qian, Zongxia Li 等CVPR 2026 · 被引用 5 次
- Cycle-Consistent Tuning for Layered Image DecompositionZheng Gu, Min Lu, Zhida Sun, Dani Lischinski 等CVPR 2026
- RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image DecompositionBinhao Wang, Shihao Zhao, Bo Cheng, Qiuyu Ji 等ICML 2026
- LayerFlow: A Unified Model for Layer-aware Video GenerationSihui Ji, Hao Luo, Xi Chen, Yuanpeng Tu 等SIGGRAPH 2025 · 被引用 12 次
- OcclusionFormer: Arranging Z-Order for Layout-Grounded Image GenerationZiye Li, Henghui DingICML 2026
