Convolutional Hierarchical Attention Network for Query-Focused Video Summarization
Shuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan, Min Yang
Abstract
Previous approaches for video summarization mainly concentrate on finding the most diverse and representative visual contents as video summary without considering the users preference. This paper addresses the task of query-focused video summarization, which takes users query and a long video as inputs and aims to generate a query-focused video summary. In this paper, we consider the task as a problem of computing similarity between video shots and query. To this end, we propose a method, named Convolutional Hierarchical Attention Network (CHAN), which consists of two parts: feature encoding network and query-relevance computing module. In the encoding network, we employ a convolutional network with local self-attention mechanism and query-aware global attention mechanism to learns visual information of each shot. The encoded features will be sent to query-relevance computing module to generate queryfocused video summary. Extensive experiments on the benchmark dataset demonstrate the competitive performance and show the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 536e83b7-acc3-4507-9f76-c1cc5bcda3b5Cited by top-tier papers9
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
- PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset SelectionSuraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff A. Bilmes et al.AAAI 2022 · 66 citations
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 36 citations
- Multiple Pairwise Ranking Networks for Personalized Video SummarizationYassir Saquil, Da Chen, Yuan He, Chuan Li et al.ICCV 2021 · 26 citations
Related papers
- Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-CaptioningYunbin Tu, Liang Li, Li Su, Qingming HuangAAAI 2025 · 1 citation
- DeepQAMVS: Query-Aware Hierarchical Pointer Networks for Multi-Video SummarizationSafa Messaoud, Ismini Lourentzou, Assma Boughoula, Mona Zehni et al.SIGIR 2021 · 12 citations
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 17 citations
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 2 citations
