Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma, Jiasong Feng, Ke Cao, Jing Wang, Yun Wang, Quanwei Zhang, Zhanjie Zhang
Abstract
Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and interframe interaction. Additionally, the scarcity of high-quality data in this field makes it difficult to precisely control storytelling tasks, including the subject's position, appearance, clothing, expression, and posture, thereby hindering further advancements. In this paper, we demonstrate that layout conditions, such as the subject's position and detailed attributes, effectively facilitate fine-grained interactions between frames. This not only strengthens the consistency of the generated frame sequence but also allows for precise control over the subject's position, appearance, and other key details. Building on this, we introduce an advanced storytelling task: Layout-Toggable Storytelling, which enables precise subject control by incorporating layout conditions. To address the lack of high-quality datasets with layout annotations for this task, we develop Lay2Story-1M, which contains over 1 million 720p and higher-resolution images, processed from approximately 11,300 hours of cartoon videos. Building on Lay2Story1M, we create Lay2Story-Bench, a benchmark with 3,000 prompts designed to evaluate the performance of different methods on this task. Furthermore, we propose Lay2Story, a robust framework based on the Diffusion Transformers (DiTs) architecture for Layout-Togglable Storytelling tasks. Through both qualitative and quantitative experiments, we find that our method outperforms the previous state-of-theart (SOTA) techniques, achieving the best results in terms of consistency, semantic correlation, and aesthetic quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 329143d6-5957-47c0-b4a0-d2eb82e30352Cited by top-tier papers6
- RelaCtrl: Relevance-Guided Efficient Control for Diffusion TransformersKe Cao, Jing Wang, Ao Ma, Jiasong Feng et al.AAAI 2026 · 15 citations
- InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster GenerationYuxin Qin, Ke Cao, Haowei Liu, Ao Ma et al.CVPR 2026 · 5 citations
- HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product ImagesYi Chen Liu, Donghao Zhou, Jie Wang, Xin Gao et al.CVPR 2026 · 5 citations
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationRun Ling, Ke Cao, Jian Lu, Ao Ma et al.AAAI 2026 · 4 citations
- Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive ModelsYexing Xu, Wei Feng, Shen Zhang, Haohan Wang et al.CVPR 2026 · 1 citation
Builds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video ModelsPatrick Kwon, Chen ChenCVPR 2026 · 1 citation
- BindWeave: Subject-Consistent Video Generation via Cross-Modal IntegrationZhaoyang Li, Dongjun Qian, Kai Su, qishuai diao et al.ICLR 2026 · 23 citations
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 19 citations
- CoDi: Subject-Consistent and Pose-Diverse Text-to-Image GenerationZhanxin Gao, Beier Zhu, Liangyao, Jian Yang et al.ICLR 2026 · 1 citation
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion TransformersSida Huang, Siqi Huang, Ping Luo, Hongyuan ZhangAAAI 2026 · 5 citations
