Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization
Litian Zhang, Xiaoming Zhang, Junshu Pan
Abstract
Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observed that different modalities of data in the news report correlate hierarchically. Traditional MSMO methods indistinguishably handle different modalities of data by learning a representation for the whole data, which is not directly adaptable to the heterogeneous contents and hierarchical correlation. In this paper, we propose a hierarchical cross-modality semantic correlation learning model (HCSCL) to learn the intra- and inter-modal correlation existing in the multimodal data. HCSCL adopts a graph network to encode the intra-modal correlation. Then, a hierarchical fusion framework is proposed to learn the hierarchical correlation between text and images. Furthermore, we construct a new dataset with relevant image annotation and image object label information to provide the supervision information for the learning procedure. Extensive experiments on the dataset show that HCSCL significantly outperforms the baseline methods in automatic summarization metrics and fine-grained diversity tests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7878207d-2e23-40f2-b354-c34f24392d63Cited by top-tier papers15
- Reinforced Adaptive Knowledge Learning for Multimodal Fake News DetectionLitian Zhang, Xiaoming Zhang, Ziyi Zhou, Feiran Huang et al.AAAI 2024 · 54 citations
- DIUSum: Dynamic Image Utilization for Multimodal SummarizationMin Xiao, Junnan Zhu, Feifei Zhai, Yu Zhou et al.AAAI 2024 · 10 citations
- Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News DetectionZiyi Zhou, Xiaoming Zhang, Shenghan Tan, Litian Zhang et al.AAAI 2025 · 9 citations
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar et al.CVPR 2024 · 7 citations
Builds on4
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang et al.EMNLP 2020 · 54 citations
- Hierarchical Scene Graph Encoder-Decoder for Image Paragraph CaptioningXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiACM MM 2020 · 25 citations
Related papers
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
- Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution ConstraintsQinfu Xu, Yiwei Wei, Chunlei Wu, Leiquan Wang et al.AAAI 2025 · 3 citations
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
- Hierarchical Semantic Enhancement Network for Multimodal Fake News DetectionQiang Zhang, Jiawei Liu, Fanrui Zhang, Jingyi Xie et al.ACM MM 2023 · 10 citations
- Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive TrainingPeng Wang, Jiangheng Wu, Xiaohang ChenSIGIR 2022 · 52 citations
