Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization
Litian Zhang, Xiaoming Zhang, Junshu Pan
摘要
Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observed that different modalities of data in the news report correlate hierarchically. Traditional MSMO methods indistinguishably handle different modalities of data by learning a representation for the whole data, which is not directly adaptable to the heterogeneous contents and hierarchical correlation. In this paper, we propose a hierarchical cross-modality semantic correlation learning model (HCSCL) to learn the intra- and inter-modal correlation existing in the multimodal data. HCSCL adopts a graph network to encode the intra-modal correlation. Then, a hierarchical fusion framework is proposed to learn the hierarchical correlation between text and images. Furthermore, we construct a new dataset with relevant image annotation and image object label information to provide the supervision information for the learning procedure. Extensive experiments on the dataset show that HCSCL significantly outperforms the baseline methods in automatic summarization metrics and fine-grained diversity tests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Reinforced Adaptive Knowledge Learning for Multimodal Fake News DetectionLitian Zhang, Xiaoming Zhang, Ziyi Zhou, Feiran Huang 等AAAI 2024 · 被引用 54 次
- DIUSum: Dynamic Image Utilization for Multimodal SummarizationMin Xiao, Junnan Zhu, Feifei Zhai, Yu Zhou 等AAAI 2024 · 被引用 10 次
- Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News DetectionZiyi Zhou, Xiaoming Zhang, Shenghan Tan, Litian Zhang 等AAAI 2025 · 被引用 9 次
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian 等CVPR 2026 · 被引用 9 次
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar 等CVPR 2024 · 被引用 7 次
它引用的顶会 Paper4
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li 等AAAI 2020 · 被引用 113 次
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan 等EMNLP 2020 · 被引用 65 次
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang 等EMNLP 2020 · 被引用 54 次
- Hierarchical Scene Graph Encoder-Decoder for Image Paragraph CaptioningXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiACM MM 2020 · 被引用 25 次
相关 Paper
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui 等CVPR 2023
- Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution ConstraintsQinfu Xu, Yiwei Wei, Chunlei Wu, Leiquan Wang 等AAAI 2025 · 被引用 3 次
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
- Hierarchical Semantic Enhancement Network for Multimodal Fake News DetectionQiang Zhang, Jiawei Liu, Fanrui Zhang, Jingyi Xie 等ACM MM 2023 · 被引用 10 次
- Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive TrainingPeng Wang, Jiangheng Wu, Xiaohang ChenSIGIR 2022 · 被引用 52 次
