Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin Jin
摘要
Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video captioning and video story generation have made some progress. However, in practical applications, we typically require synchronized narrations for ongoing visual scenes. In this work, we introduce a new task of Synchronized Video Storytelling, which aims to generate synchronous and informative narrations for videos. These narrations, associated with each video clip, should relate to the visual content, integrate relevant knowledge, and have an appropriate word count corresponding to the clip's duration. Specifically, a structured storyline is beneficial to guide the generation process, ensuring coherence and integrity. To support the exploration of this task, we introduce a new benchmark dataset E-SyncVidStory with rich annotations. Since existing Multimodal LLMs are not effective in addressing this task in oneshot or few-shot settings, we propose a framework named VideoNarrator that can generate a storyline for input videos and simultaneously generate narrations with the guidance of the generated or predefined storyline. We further introduce a set of evaluation metrics to thoroughly assess the generation. Both automatic and human evaluations validate the effectiveness of our approach. Our dataset, codes, and evaluations will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan 等ICLR 2026 · 被引用 215 次
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein 等NeurIPS 2025 · 被引用 132 次
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo 等ICLR 2026 · 被引用 92 次
- Stable Video Infinity: Infinite-Length Video Generation with Error RecyclingWuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao 等ICLR 2026 · 被引用 69 次
- Long Context Tuning for Video GenerationYuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingYaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan 等CVPR 2022 · 被引用 31 次
相关 Paper
- HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External KnowledgeXueyan Wang, Dingyi Yang, Qin JinACL 2026
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
- Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot VideosMingfei Han, Linjie Yang, Xiaojun Chang, Lina Yao 等ICLR 2025 · 被引用 3 次
- Learning Video Representations from Large Language ModelsYue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit GirdharCVPR 2023
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic 等EMNLP 2022 · 被引用 2 次
