Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin Jin
Abstract
Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video captioning and video story generation have made some progress. However, in practical applications, we typically require synchronized narrations for ongoing visual scenes. In this work, we introduce a new task of Synchronized Video Storytelling, which aims to generate synchronous and informative narrations for videos. These narrations, associated with each video clip, should relate to the visual content, integrate relevant knowledge, and have an appropriate word count corresponding to the clip's duration. Specifically, a structured storyline is beneficial to guide the generation process, ensuring coherence and integrity. To support the exploration of this task, we introduce a new benchmark dataset E-SyncVidStory with rich annotations. Since existing Multimodal LLMs are not effective in addressing this task in oneshot or few-shot settings, we propose a framework named VideoNarrator that can generate a storyline for input videos and simultaneously generate narrations with the guidance of the generated or predefined storyline. We further introduce a set of evaluation metrics to thoroughly assess the generation. Both automatic and human evaluations validate the effectiveness of our approach. Our dataset, codes, and evaluations will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c1315fc-6a07-4256-9566-890eb27cfbd8Cited by top-tier papers10
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan et al.ICLR 2026 · 215 citations
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein et al.NeurIPS 2025 · 132 citations
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo et al.ICLR 2026 · 92 citations
- Stable Video Infinity: Infinite-Length Video Generation with Error RecyclingWuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao et al.ICLR 2026 · 69 citations
- Long Context Tuning for Video GenerationYuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma et al.ICCV 2025 · 6 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingYaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan et al.CVPR 2022 · 31 citations
Related papers
- HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External KnowledgeXueyan Wang, Dingyi Yang, Qin JinACL 2026
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot VideosMingfei Han, Linjie Yang, Xiaojun Chang, Lina Yao et al.ICLR 2025 · 3 citations
- Learning Video Representations from Large Language ModelsYue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit GirdharCVPR 2023
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic et al.EMNLP 2022 · 2 citations
