Multimodal Summarization with Guidance of Multimodal Reference
Junnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li, Chengqing Zong, Changliang Li
Abstract
Multimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Orderranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
- DOC2PPT: Automatic Presentation Slides Generation from Scientific DocumentsTsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale SongAAAI 2022 · 83 citations
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang et al.EMNLP 2020 · 54 citations
Related papers
- Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal SummarizationLitian Zhang, Xiaoming Zhang, Junshu PanAAAI 2022 · 53 citations
- ModalSyncSum: Synchronizing Image and Text for Reliable Summary GenerationXuanqi Chen, Ziying Rong, Xinfeng Liao, Yiqian Wu et al.AAAI 2026
- DIUSum: Dynamic Image Utilization for Multimodal SummarizationMin Xiao, Junnan Zhu, Feifei Zhai, Yu Zhou et al.AAAI 2024 · 10 citations
- Multi-Modal Supplementary-Complementary Summarization using Multi-Objective OptimizationAnubhav Jangra, Sriparna Saha, Adam Jatowt, Mohammed HasanuzzamanSIGIR 2021 · 18 citations
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
