DeepQAMVS: Query-Aware Hierarchical Pointer Networks for Multi-Video Summarization
Safa Messaoud, Ismini Lourentzou, Assma Boughoula, Mona Zehni, Zhizhen Zhao, Chengxiang Zhai, Alexander G. Schwing
Abstract
The recent growth of web video sharing platforms has increased the demand for systems that can efficiently browse, retrieve and summarize video content. Query-aware multi-video summarization is a promising technique that caters to this demand. In this work, we introduce a novel Query-Aware Hierarchical Pointer Network for Multi-Video Summarization, termed DeepQAMVS, that jointly optimizes multiple criteria: (1) conciseness, (2) representativeness of important query-relevant events and (3) chronological soundness. We design a hierarchical attention model that factorizes over three distributions, each collecting evidence from a different modality, followed by a pointer network that selects frames to include in the summary. DeepQAMVS is trained with reinforcement learning, incorporating rewards that capture representativeness, diversity, query-adaptability and temporal coherence. We achieve state-of-the-art results on the MVS1K dataset, with inference time scaling linearly with the number of input video frames.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fe9e1dc-1e4c-4995-9ffd-6bad0418ce79Cited by top-tier papers4
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 36 citations
- AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationMilton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen et al.CVPR 2026 · 4 citations
- Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video CaptioningYolo Yunlong Tang, Chao Huang, Susan Liang, Jing Bi et al.CVPR 2026 · 2 citations
Builds on1
Related papers
- Convolutional Hierarchical Attention Network for Query-Focused Video SummarizationShuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan et al.AAAI 2020 · 2 citations
- Be Relevant, Non-Redundant, and Timely: Deep Reinforcement Learning for Real-Time Event SummarizationMin Yang, Chengming Li, Fei Sun, Zhou Zhao et al.AAAI 2020 · 8 citations
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 1 citation
- Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-CaptioningYunbin Tu, Liang Li, Li Su, Qingming HuangAAAI 2025 · 1 citation
- TripleSumm: Adaptive Triple-Modality Fusion for Video SummarizationSumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim et al.ICLR 2026 · 2 citations
