CLIP-It! Language-Guided Video Summarization
Medhini Narasimhan, Anna Rohrbach, Trevor Darrell
Abstract
A generic video summary is an abridged version of a video that conveys the whole story and features the most important scenes. Yet the importance of scenes in a video is often subjective, and users should have the option of customizing the summary by using natural language to specify what is important to them. Further, existing models for fully automatic generic summarization have not exploited available language models, which can serve as an effective prior for saliency. This work introduces CLIP-It, a single framework for addressing both generic and query-focused video summarization, typically approached separately in the literature. We propose a language-guided multimodal transformer that learns to score frames in a video based on their importance relative to one another and their correlation with a user-defined query (for query-focused summarization) or an automatically generated dense video caption (for generic video summarization). Our model can be extended to the unsupervised setting by training without ground-truth supervision. We outperform baselines and prior work by a significant margin on both standard video summarization datasets (TVSum and SumMe) and a query-focused video summarization dataset (QFVS). Particularly, we achieve large improvements in the transfer setting, attesting to our method's strong generalization capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fca2898-7daf-457d-a389-e5a0aecea904Cited by top-tier papers33
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIPAndreas Fürst, Elisabeth Rumetshofer, Johannes Lehner, Viet T. Tran et al.NeurIPS 2022 · 131 citations
- Language-Bridged Spatial-Temporal Interaction for Referring Video Object SegmentationZihan Ding, Tianrui Hui, Junshi Huang, Xiaoming Wei et al.CVPR 2022 · 62 citations
- V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction TuningHang Hua, Yunlong Tang, Chenliang Xu, Jiebo LuoAAAI 2025 · 61 citations
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang et al.NeurIPS 2023 · 60 citations
Builds on2
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu et al.ACL 2020 · 168 citations
Related papers
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 1 citation
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 2 citations
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 36 citations
