Masked Region Transformer for Layered Image Generation and Editing at Scale
Zhicong Tang, Jingye Chen, Zhao Zhang, Mohan Zhou, Yuchi Liu, Yifan Pu, Yalong Bai, Ethan Smith, Yuhui Yuan
Abstract
Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of the generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To address this gap, we present the Masked Region Transformer, a 20B-parameter diffusion model tailored for multi-layer transparent image generation and editing, trained on over 10M multilingual design samples spanning diverse aspect ratios and textual prompts. To fully leverage this scale, we make three key technical contributions. First, we unify three complementary tasks---text-to-layers, image-to-layers, and layers-to-layers---within a shared masked region diffusion framework, where selective token masking enables flexible cross-modal generation and fine-grained layer-wise editing. Second, we design an efficient conditional diffusion decoder that incorporates Gated DeltaNet and gated attention mechanisms, enhancing visual fidelity while maintaining computational efficiency. Third, we introduce an overflow-aware canvas layer to handle boundary inconsistencies and support semi-transparent background synthesis, enabling complete editable layer generation beyond visible canvas boundaries. Additionally, we apply distribution matching distillation to achieve one-step, real-time multi-layer generation with minimal quality degradation. Extensive experiments demonstrate that our framework substantially outperforms prior state-of-the-art approaches across all three tasks, establishing a new benchmark for region-aware transparent image generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e0b66eb-4a49-40bb-9774-080c03ba2759Builds on39
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
Related papers
- ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image GenerationYifan Pu, Yiming Zhao, Zhicong Tang, Ruihong Yin et al.CVPR 2025
- DreamLayer: Simultaneous Multi-Layer Generation via Diffusion ModelJunjia Huang, Pengxiang Yan, Jinhang Cai, Jiyang Liu et al.ICCV 2025 · 4 citations
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li et al.CVPR 2026 · 14 citations
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and GenerationShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.ICLR 2026 · 14 citations
- Qwen-Image-Layered: Towards Inherent Editability via Layer DecompositionShengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao et al.CVPR 2026 · 30 citations
