Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, Zeyu Zheng, Zirui Wang, Cihang Xie, Yuyin Zhou
Abstract
This paper introduces Story-Iter, a new training-free iterative paradigm to enhance long-story generation. Unlike existing methods that rely on fixed reference images to construct a complete story, our approach features a novel external iterative paradigm, extending beyond the internal iterative denoising steps of diffusion models, to continuously refine each generated image by incorporating all reference images from the previous round. To achieve this, we propose a plug-and-play, training-free global reference cross-attention (GRCA) module, modeling all reference frames with global embeddings, ensuring semantic consistency in long sequences. By progressively incorporating holistic visual context and text constraints, our iterative paradigm enables precise generation with fine-grained interactions, optimizing the story visualization step-by-step. Extensive experiments in the official story visualization dataset and our long story benchmark demonstrate that Story-Iter's state-of-the-art performance in long-story visualization (up to 100 frames) excels in both semantic consistency and fine-grained interactions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 591702df-326a-479a-b9b8-bde534fc82c9Cited by top-tier papers4
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang et al.ICCV 2025 · 13 citations
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen et al.NeurIPS 2025 · 9 citations
- IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual PromptingYuxin Zhang, Minyan Luo, Weiming Dong, Xiao Yang et al.SIGGRAPH 2025 · 2 citations
- DreamShot: Personalized Storyboard Synthesis with Video Diffusion PriorJunjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu et al.CVPR 2026 · 1 citation
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
Related papers
- Consistent Story Generation: Unlocking the Potential of Zigzag SamplingMingxiao Li, Mang Ning, Marie-Francine MoensNeurIPS 2025 · 2 citations
- Infinite-Story: A Training-Free Consistent Text-to-Image GenerationJihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo et al.AAAI 2026 · 1 citation
- Storybooth: Training-Free Multi-Subject Consistency for Improved Visual StorytellingJaskirat Singh, Junshen K. Chen, Jonas Kohler, Michael F. CohenICLR 2025
- FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal AttentionYu Lu, Yuanzhi Liang, Linchao Zhu, Yi YangNeurIPS 2024 · 101 citations
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
