SQuALITY: Building a Long-Document Summarization Dataset the Hard Way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, Samuel R. Bowman
摘要
Summarization datasets are often assembled either by scraping naturally occurring publicdomain summaries-which are nearly always in difcult-to-work-with technical domainsor by using approximate heuristics to extract them from everyday text-which frequently yields unfaithful summaries. In this work, we turn to a slower but more straightforward approach to developing summarization benchmark data: We hire highly-qualied contractors to read stories and write original summaries from scratch. To amortize reading time, we collect ve summaries per document, with the rst giving an overview and the subsequent four addressing specic questions. We use this protocol to collect SQuAL-ITY, a dataset of question-focused summaries built on the same public-domain short stories as the multiple-choice dataset QuALITY (Pang et al., 2021b). Experiments with stateof-the-art summarization systems show that our dataset is challenging and that existing automatic evaluation metrics are weak indicators of quality. SQuALITY is available at https: //github.com/nyu-mll/SQuALITY .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- Block-Recurrent TransformersDeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer 等NeurIPS 2022 · 被引用 163 次
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang 等ICML 2023 · 被引用 76 次
- PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational PathsBoyu Chen, Zirui Guo, Zidan Yang, Yuluo Chen 等AAAI 2026 · 被引用 45 次
- LooGLE: Can Long-Context Language Models Understand Long Contexts?Jiaqi Li, Mengmeng Wang, Zilong Zheng, Muhan ZhangACL 2024 · 被引用 32 次
它引用的顶会 Paper8
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 被引用 90 次
- SCROLLS: Standardized CompaRison Over Long Language SequencesUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat 等EMNLP 2022 · 被引用 37 次
相关 Paper
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 被引用 52 次
- Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationArtidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng WuACL 2023 · 被引用 4 次
- QuerySum: A Multi-Document Query-Focused Summarization Dataset Augmented with Similar Query ClustersYushan Liu, Zili Wang, Ruifeng YuanAAAI 2024 · 被引用 14 次
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams 等EMNLP 2024 · 被引用 2 次
- SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationElizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez 等EMNLP 2023 · 被引用 6 次
