CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization
Min Xiao, Junnan Zhu, Haitao Lin, Yu Zhou, Chengqing Zong
Abstract
Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear. Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are useful. Therefore, we propose a novel Coarse-to-Fine contribution network for multimodal Summarization (CFSum) to consider different contributions of images for summarization. First, to eliminate the interference of useless images, we propose a pre-filter module to abandon useless images. Second, to make accurate use of useful images, we propose two levels of visual complement modules, word level and phrase level. Specifically, image contributions are calculated and are adopted to guide the attention of both textual and visual modalities. Experimental results have shown that CFSum significantly outperforms multiple strong baselines on the standard benchmark. Furthermore, the analysis verifies that useful images can even help generate nonvisual words which are implicitly represented in the image 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72ea7ad0-5ec3-47b8-8b63-28f6a57cb431Cited by top-tier papers4
- DIUSum: Dynamic Image Utilization for Multimodal SummarizationMin Xiao, Junnan Zhu, Feifei Zhai, Yu Zhou et al.AAAI 2024 · 10 citations
- Exploring the Trade-Off within Visual Information for MultiModal Sentence SummarizationMinghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang et al.SIGIR 2024 · 3 citations
- SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yining Zhang, Yupu Liang et al.EMNLP 2025
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang et al.ACL 2025
Builds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Aspect-Aware Multimodal Summarization for Chinese E-Commerce ProductsHaoran Li, Peng Yuan, Song Xu, Youzheng Wu et al.AAAI 2020 · 79 citations
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive SummarizationTiezheng Yu, Wenliang Dai, Zihan Liu, Pascale FungEMNLP 2021 · 64 citations
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang et al.EMNLP 2020 · 54 citations
Related papers
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie et al.ACM MM 2020 · 159 citations
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal SummarizationLitian Zhang, Xiaoming Zhang, Junshu PanAAAI 2022 · 53 citations
