V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
Hang Hua, Yunlong Tang, Chenliang Xu, Jiebo Luo
Abstract
Video summarization aims to create short, accurate, and cohesive summaries of longer videos. Despite the existence of various video summarization datasets, a notable limitation is their limited amount of source videos, which hampers the effective training of advanced large vision-language models (VLMs). Additionally, most existing datasets are created for video-to-video summarization, overlooking the contemporary need for multimodal video content summarization. Recent efforts have been made to expand from unimodal to multimodal video summarization, categorizing the task into three sub-tasks based on the summary's modality: video-to-video (V2V), video-to-text (V2T), and a combination of video and text summarization (V2VT). However, the textual summaries in previous multimodal datasets are inadequate. To address these issues, we introduce Instruct-V2Xum, a cross-modal video summarization dataset featuring 30,000 diverse videos sourced from YouTube, with lengths ranging from 40 to 940 seconds and an average summarization ratio of 16.39%. Each video summary in Instruct-V2Xum is paired with a textual summary that references specific frame indexes, facilitating the generation of aligned video and textual summaries. In addition, we propose a new video summarization framework named V2Xum-LLM. V2Xum-LLM, specifically V2Xum-LLaMA in this study, is the first framework that unifies different video summarization tasks into one large language model's (LLM) text decoder and achieves task-controllable video summarization with temporal prompts and task instructions. Experiments show that V2Xum-LLaMA outperforms strong baseline models on multiple video summarization tasks. Furthermore, we propose an enhanced evaluation metric for V2V and V2VT summarization tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- PromptFix: You Prompt and We Fix the PhotoYongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu et al.NeurIPS 2024 · 55 citations
- UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban SpacesBaining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang et al.ACL 2025 · 31 citations
- Latent Chain-of-Thought for Visual ReasoningGuohao Sun, Hang Hua, Jian Wang, Jiebo Luo et al.NeurIPS 2025 · 30 citations
- VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingYongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng et al.AAAI 2025 · 27 citations
- Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingJingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang et al.NeurIPS 2025 · 25 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 2 citations
- Scaling Up Video Summarization Pretraining with Large Language ModelsDawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron, Hanieh Deilamsalehy et al.CVPR 2024
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar et al.CVPR 2024 · 7 citations
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
