Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement Learning
Yuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren, Jiawei Han
Abstract
While neural sequence learning methods have made significant progress in single-document summarization (SDS), they produce unsatisfactory results on multi-document summarization (MDS). We observe two major challenges when adapting SDS advances to MDS: (1) MDS involves larger search space and yet more limited training data, setting obstacles for neural methods to learn adequate representations; (2) MDS needs to resolve higher information redundancy among the source documents, which SDS methods are less effective to handle. To close the gap, we present RL-MMR, Maximal Margin Relevance-guided Reinforcement Learning for MDS, which unifies advanced neural SDS methods and statistical measures used in classical MDS. RL-MMR casts MMR guidance on fewer promising candidates, which restrains the search space and thus leads to better representation learning. Additionally, the explicit redundancy measure in MMR helps the neural representation of the summary to better capture redundancy. Extensive experiments demonstrate that RL-MMR achieves state-of-the-art performance on benchmark MDS datasets. In particular, we show the benefits of incorporating MMR into end-to-end learning when adapting SDS to MDS in terms of both learning effectiveness and efficiency. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d88a1b61-f3db-498d-9c89-bde8e59b20ebCited by top-tier papers10
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 147 citations
- A Multi-Document Coverage Reward for RELAXed Multi-Document SummarizationJacob Parnell, Inigo Jauregi Unanue, Massimo PiccardiACL 2022 · 16 citations
- MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingGabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor et al.ACL 2025 · 13 citations
- PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets StreamSusik Yoon, Hou Pong Chan, Jiawei HanWWW 2023 · 13 citations
- CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited SupervisionYuning Mao, Ming Zhong, Jiawei HanEMNLP 2022 · 11 citations
Builds on4
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Generating Representative Headlines for News StoriesXiaotao Gu, Yuning Mao, Jiawei Han, Jialu Liu et al.WWW 2020 · 77 citations
- Facet-Aware Evaluation for Extractive SummarizationYuning Mao, Liyuan Liu, Qi Zhu, Xiang Ren et al.ACL 2020 · 19 citations
Related papers
- SgSum: Transforming Multi-document Summarization into Sub-graph SelectionMoye Chen, Wei Li, Jiachen Liu, Xinyan Xiao et al.EMNLP 2021 · 21 citations
- Content- and Topology-Aware Representation Learning for Scientific Multi-LiteratureKai Zhang, Kaisong Song, Yangyang Kang, Xiaozhong LiuEMNLP 2023
- A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document SummarizationShiyin Tan, Jaeeon Park, Dongyuan Li, Renhe Jiang et al.SIGIR 2025
- Leveraging Graph to Improve Abstractive Multi-Document SummarizationWei Li, Xinyan Xiao, Jiachen Liu, Hua Wu et al.ACL 2020 · 118 citations
- Be Relevant, Non-Redundant, and Timely: Deep Reinforcement Learning for Real-Time Event SummarizationMin Yang, Chengming Li, Fei Sun, Zhou Zhao et al.AAAI 2020 · 8 citations
