Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior Aesthetics
Keming Wu, Junwen Chen, Zhanhao Liang, Yinuo Wang, Ji Li, Chao Zhang, Bin Wang, Yuhui Yuan
摘要
first fine-tunes the DiTs (e.g., SD3) to follow an anonymous layout, then continues fine-tuning the DiTs to follow the semantic layout, and finally includes a quality-tuning stage to enhance visual aesthetics. We show that this hybrid design is highly data-efficient, as we find only using a small amount of semantic layout annotations is sufficient, thereby significantly reducing dependency on regional prompts. In addition, we propose an efficient regional diffusion transformer to encode the spatial layout information using just a set of lower-resolution regional tokens instead of various carefully designed layout tokens. The region-wise diffusion loss over these regional tokens can guide the diffusion transformer learn to follow the given layout implicitly. We empirically validate the effectiveness of our approach by comparing it with the latest version of SiamLayout and show that our method achieves better results while being more than 10× more data efficient and ensuring superior aesthetics. Project Page: https://hybrid-layout-msra.github.io Attention Ratio: 80.86% Area Ratio: 67.72% Attention Ratio: 82.31% Area Ratio: 72.90% Attention Ratio: 79.40% Area Ratio: 67.19% Attention Ratio: 68.95% Area Ratio: 60.21% Attention Ratio: 81.47% Area Ratio: 76.22%
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion TransformersRuidong Chen, Yancheng Bai, Xuanpu Zhang, Jianhao Zeng 等CVPR 2026 · 被引用 9 次
- Comp-Attn: Present-and-Align Attention for Compositional Video GenerationHongyu Zhang, Yufan Deng, Shenghai Yuan, Xuehan Hou 等ICML 2026
它引用的顶会 Paper30
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
相关 Paper
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao 等ICCV 2025 · 被引用 7 次
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick 等ICCV 2025 · 被引用 9 次
- Dynamic Diffusion TransformerWangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 等ICLR 2025
- DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and PerceptionYibo Wang, Ruiyuan Gao, Kai Chen, Kaiqiang Zhou 等CVPR 2024 · 被引用 14 次
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang 等ICCV 2025 · 被引用 13 次
