QuerySum: A Multi-Document Query-Focused Summarization Dataset Augmented with Similar Query Clusters
Yushan Liu, Zili Wang, Ruifeng Yuan
Abstract
Query-focused summarization (QFS) aims to summarize the source document(s) with regard to a specific aspect of information given in a query. It plays an important role in presenting users with a concise answer summary from a set of query-relevant documents retrieved by the information retrieval system. Nonetheless, the QFS research has long been hampered by the lack of adequate datasets in terms of both quality and quantity. In this paper, we introduce a large-scale multi-document query-focused summarization dataset, called QuerySum, which contains 27,041 data samples covering diverse topics and its quality is guaranteed through human verification. Unlike some previous QFS datasets constructed directly from the question answering datasets, 74% queries in our dataset are the challenging non-factoid What-, Why-, and How-questions. More importantly, we also provide a set of similar queries together with the corresponding summaries pairs for each query as the retrieved context, presenting a new feature of QuerySum. We aim to encourage research efforts in query intention understanding in the context of QFS. Leveraging QuerySum's depth, we propose a model for query-aware multi-document summarization and set a new QFS benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92d698ae-22c8-49ef-a4e1-7c28c25be194Cited by top-tier papers4
- Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language ModelsJunjie Wu, Gefei Gu, Yanan Zheng, Dit-Yan Yeung et al.ACL 2025 · 4 citations
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length ContextsYuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min et al.EMNLP 2025
- WikiREVIEW: A Multi-Perspective Review Framework for Automatic Wiki-Style Article GenerationGuo-Biao Zhang, Zhijing Wu, Tian Lan, Ding-Yuan Liu et al.AAAI 2026
- Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from DocumentsAnkan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick et al.EMNLP 2025
Builds on3
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani et al.EMNLP 2022 · 59 citations
- Multi-hop Inference for Question-driven SummarizationYang Deng, Wenxuan Zhang, Wai LamEMNLP 2020 · 17 citations
Related papers
- Few-shot Query-Focused Summarization with Prefix-MergingRuifeng Yuan, Zili Wang, Ziqiang Cao, Wenjie LiEMNLP 2022 · 6 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- Data Augmentation for Abstractive Query-Focused Multi-Document SummarizationRamakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong et al.AAAI 2021 · 46 citations
- OpenAsp: A Benchmark for Multi-document Open Aspect-based SummarizationShmuel Amar, Liat Schiff, Ori Ernst, Asi Shefer et al.EMNLP 2023 · 3 citations
- QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question AnsweringAn Quang Tang, Xiuzhen Zhang, Minh Ngoc Dinh, Zhuang LiACL 2025
