A Topic-aware Summarization Framework with Different Modal Side Information
Xiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng, Qiang Yang, Qishen Zhang, Xin Gao, Xiangliang Zhang
Abstract
Automatic summarization plays an important role in the exponential document growth on the Web. On content websites such as CNN.com and WikiHow.com, there often exist various kinds of side information along with the main document for attention attraction and easier understanding, such as videos, images, and queries. Such information can be used for better summarization, as they often explicitly or implicitly mention the essence of the article. However, most of the existing side-aware summarization methods are designed to incorporate either single-modal or multi-modal side information, and cannot effectively adapt to each other. In this paper, we propose a general summarization framework, which can flexibly incorporate various modalities of side information. The main challenges in designing a flexible summarization model with side information include: (1) the side information can be in textual or visualformat, and the model needs to align and unify it with the document into the same semantic space, (2) the side inputs can contain information from variousaspects, and the model should recognize the aspects useful for summarization. To address these two challenges, we first propose a unified topic encoder, which jointly discovers latent topics from the document and various kinds of side information. The learned topics flexibly bridge and guide the information flow between multiple inputs in a graph encoder through a topic-aware interaction. We secondly propose a triplet contrastive learning mechanism to align the single-modal or multi-modal information into a unified semantic space, where thesummary quality is enhanced by better understanding thedocument andside information. Results show that our model significantly surpasses strong baselines on three public single-modal or multi-modal benchmark summarization datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
- The Truth Becomes Clearer Through Debate! Multi-Agent Systems with Large Language Models Unmask Fake NewsYuhan Liu, Yuxuan Liu, Xiaoqing Zhang, Xiuying Chen et al.SIGIR 2025 · 20 citations
- P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday TaskWeiye Xu, Min Wang, Wengang Zhou, Houqiang LiACM MM 2024 · 5 citations
- The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM AgentsYuhan Liu, Zirui Song, Juntian Zhang, Xiaoqing Zhang et al.EMNLP 2025 · 2 citations
Builds on15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 117 citations
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
Related papers
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 1 citation
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
- A Knowledge Augmented and Multimodal-Based Framework for Video SummarizationJiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu YangACM MM 2022 · 11 citations
- ModalSyncSum: Synchronizing Image and Text for Reliable Summary GenerationXuanqi Chen, Ziying Rong, Xinfeng Liao, Yiqian Wu et al.AAAI 2026
