Convolutional Hierarchical Attention Network for Query-Focused Video Summarization
Shuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan, Min Yang
摘要
Previous approaches for video summarization mainly concentrate on finding the most diverse and representative visual contents as video summary without considering the users preference. This paper addresses the task of query-focused video summarization, which takes users query and a long video as inputs and aims to generate a query-focused video summary. In this paper, we consider the task as a problem of computing similarity between video shots and query. To this end, we propose a method, named Convolutional Hierarchical Attention Network (CHAN), which consists of two parts: feature encoding network and query-relevance computing module. In the encoding network, we employ a convolutional network with local self-attention mechanism and query-aware global attention mechanism to learns visual information of each shot. The encoded features will be sent to query-relevance computing module to generate queryfocused video summary. Extensive experiments on the benchmark dataset demonstrate the competitive performance and show the effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick 等ICCV 2023 · 被引用 221 次
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 等ICCV 2023 · 被引用 152 次
- PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset SelectionSuraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff A. Bilmes 等AAAI 2022 · 被引用 66 次
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 被引用 36 次
- Multiple Pairwise Ranking Networks for Personalized Video SummarizationYassir Saquil, Da Chen, Yuan He, Chuan Li 等ICCV 2021 · 被引用 26 次
相关 Paper
- Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-CaptioningYunbin Tu, Liang Li, Li Su, Qingming HuangAAAI 2025 · 被引用 1 次
- DeepQAMVS: Query-Aware Hierarchical Pointer Networks for Multi-Video SummarizationSafa Messaoud, Ismini Lourentzou, Assma Boughoula, Mona Zehni 等SIGIR 2021 · 被引用 12 次
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 被引用 17 次
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 被引用 196 次
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 被引用 2 次
