Multimodal Video Summarization via Time-Aware Transformers
Xindi Shang, Zehuan Yuan, Anran Wang, Changhu Wang
摘要
With the growing number of videos in video sharing platforms, how to facilitate the searching and browsing of the user-generated video has attracted intense attention by multimedia community. To help people efficiently search and browse relevant videos, summaries of videos become important. The prior works in multimodal video summarization mainly explore visual and ASR tokens as two separate sources and struggle to fuse the multimodal information for generating the summaries. However, the time information inside videos is commonly ignored. In this paper, we find that it is important to leverage the timestamps to accurately incorporate multimodal signals for the task. We propose a Time-Aware Multimodal Transformer (TAMT) with a novel short-term order-sensitive attention mechanism. The attention mechanism can attend the inputs differently based on time difference to explore the time information inherent inside video more thoroughly. As such, TAMT can fuse the different modalities better for summarizing the videos. Experiments show that our proposed approach is effective and achieves the state-of-the-art performances on both YouCookII and open-domain How2 datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosNayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu 等EMNLP 2022 · 被引用 10 次
- Combining Vision and Language Representations for Patch-based Identification of Lexico-Semantic RelationsPrince Jha, Gaël Dias, Alexis Lechervy, José G. Moreno 等ACM MM 2022 · 被引用 5 次
- EAGLE: Egocentric AGgregated Language-video EngineJing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi 等ACM MM 2024 · 被引用 3 次
- What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsDongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon 等ACL 2025
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang 等ACL 2025
相关 Paper
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui 等CVPR 2023
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationShilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen 等AAAI 2024 · 被引用 67 次
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 被引用 196 次
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen 等CVPR 2022 · 被引用 150 次
- Attention is not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence FusionTao Liang, Guosheng Lin, Lei Feng, Yan Zhang 等ICCV 2021 · 被引用 84 次
