Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
Zhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen, Peng Jiang, Yanye Lu
Abstract
We present Narrative Weaver, a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual content generation. While existing models excel at generating high-fidelity short-form visual content, they struggle to maintain narrative coherence and visual consistency across extended sequences-a critical limitation for realworld applications such as filmmaking and e-commerce advertising. Narrative Weaver introduces the first holistic solution that seamlessly integrates three essential capabilities: fine-grained control, automatic narrative planning, and long-range coherence. Our architecture combines a Multimodal Large Language Model (MLLM) for high-level narrative planning with a novel fine-grained control module featuring a dynamic Memory Bank that prevents visual drift. To enable practical deployment, we develop a progressive, multi-stage training strategy that efficiently leverages existing pre-trained models, achieving state-of-the-art performance even with limited training data. Recognizing the absence of suitable evaluation benchmarks, we construct and release the E-commerce Advertising Video Storyboard Dataset (EAVSD)-the first comprehensive dataset for this task, containing over 330K high-quality images with rich narrative annotations. Through extensive experiments across three distinct scenarios (controllable multi-scene generation, autonomous storytelling, and e-commerce advertising), we demonstrate our method's superiority while opening new possibilities for AI-driven content creation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90770f59-6eec-49ca-813f-2c206d3e5e96Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
Related papers
- CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationWei Chen, Lin Li, Yongqi Yang, Bin Wen et al.CVPR 2025
- DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion AdaptationZun Wang, Jialu Li, Han Lin, Jaehong Yoon et al.AAAI 2026 · 5 citations
- FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive DiffusionXiangyang Luo, Qingyu Li, Xiaokun Liu, Wenyu Qin et al.AAAI 2026 · 3 citations
- Wan-Weaver: Interleaved Multi-modal Generation via Decoupled TrainingJinbo Xing, Zeyinzi Jiang, Yuxiang Tuo, Chaojie Mao et al.CVPR 2026 · 2 citations
- MovieDreamer: Hierarchical Generation for Coherent Long Visual SequencesCanyu Zhao, Mingyu Liu, Wen Wang, Weihua Chen et al.ICLR 2025
