Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong, Yizhe Zhang, Mohit Bansal, Jianfeng Gao
Abstract
The progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two data augmentation methods: (1) transferring the commonly used single-document CNN/Daily Mail summarization dataset to create the QMDSCNN dataset, and (2) mining search-query logs to create the QMDSIR dataset. These two datasets have complementary properties, i.e., QMDSCNN has real summaries but queries are simulated, while QMDSIR has real queries but simulated summaries. To cover both these real summary and query aspects, we build abstractive end-to-end neural network models on the combined datasets that yield new state-of-the-art transfer results on DUC datasets. We also introduce new hierarchical encoders that enable a more efficient encoding of the query together with multiple documents. Empirical results demonstrate that our data augmentation and encoding methods outperform baseline models on automatic metrics, as well as on human evaluations along multiple attributes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Z-Code++: A Pre-trained Language Model Optimized for Abstractive SummarizationPengcheng He, Baolin Peng, Song Wang, Yang Liu et al.ACL 2023 · 27 citations
- Towards Explainable Search Results: A Listwise Explanation GeneratorPuxuan Yu, Razieh Rahimi, James AllanSIGIR 2022 · 26 citations
- Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading ComprehensionNaoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian et al.EMNLP 2021 · 14 citations
- MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingGabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor et al.ACL 2025 · 13 citations
- WikiHowQA: A Comprehensive Benchmark for Multi-Document Non-Factoid Question AnsweringValeria Bolotova-Baranova, Vladislav Blinov, Sofya Filippova, Falk Scholer et al.ACL 2023 · 12 citations
Related papers
- QuerySum: A Multi-Document Query-Focused Summarization Dataset Augmented with Similar Query ClustersYushan Liu, Zili Wang, Ruifeng YuanAAAI 2024 · 14 citations
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei et al.EMNLP 2020 · 42 citations
- Leveraging Graph to Improve Abstractive Multi-Document SummarizationWei Li, Xinyan Xiao, Jiachen Liu, Hua Wu et al.ACL 2020 · 118 citations
- Learning to Generate Overlap Summaries through Noisy Synthetic DataNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 1 citation
- Peek Across: Improving Multi-Document Modeling via Cross-Document Question-AnsweringAvi Caciularu, Matthew E. Peters, Jacob Goldberger, Ido Dagan et al.ACL 2023 · 8 citations
