VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles
Mingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan, Dongyan Zhao, Rui Yan
Abstract
A popular multimedia news format nowadays is providing users with a lively video and a corresponding news article, which is employed by influential news media including CNN, BBC, and social media including Twitter and Weibo. In such a case, automatically choosing a proper cover frame of the video and generating an appropriate textual summary of the article can help editors save time, and readers make the decision more effectively. Hence, in this paper, we propose the task of Videobased Multimodal Summarization with Multimodal Output (VMSMO) to tackle such a problem. The main challenge in this task is to jointly model the temporal dependency of video with semantic meaning of article. To this end, we propose a Dual-Interaction-based Multimodal Summarizer (DIMS), consisting of a dual interaction module and multimodal generator. In the dual interaction module, we propose a conditional self-attention mechanism that captures local semantic information within video and a global-attention mechanism that handles the semantic relationship between news text and video from a high level. Extensive experiments conducted on a large-scale real-world VMSMO dataset 1 show that DIMS achieves the state-of-the-art performance in terms of both automatic metrics and human evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- DOC2PPT: Automatic Presentation Slides Generation from Scientific DocumentsTsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale SongAAAI 2022 · 83 citations
- Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal SummarizationLitian Zhang, Xiaoming Zhang, Junshu PanAAAI 2022 · 53 citations
- The Style-Content Duality of Attractiveness: Learning to Write Eye-Catching Headlines via DisentanglementMingzhe Li, Xiuying Chen, Min Yang, Shen Gao et al.AAAI 2021 · 21 citations
- StreamHover: Livestream Transcript Summarization and AnnotationSangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui et al.EMNLP 2021 · 18 citations
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
Builds on2
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- Learning to Respond with Stickers: A Framework of Unifying Multi-Modality in Multi-Turn DialogShen Gao, Xiuying Chen, Chang Liu, Li Liu et al.WWW 2020 · 42 citations
Related papers
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
- TopicCAT: Unsupervised Topic-Guided Co-Attention Transformer for Extreme Multimodal SummarisationPeggy Tang, Kun Hu, Lei Zhang, Junbin Gao et al.ACM MM 2023 · 5 citations
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 2 citations
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 17 citations
- A Topic-aware Summarization Framework with Different Modal Side InformationXiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng et al.SIGIR 2023 · 10 citations
