STORYSUMM: Evaluating Faithfulness in Story Summarization
Melanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams, Lydia B. Chilton, Kathleen R. McKeown
摘要
Human evaluation has been the gold standard for checking faithfulness in abstractive summarization. However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out. We therefore introduce a new dataset, STORYSUMM, comprising LLM summaries of short stories with localized faithfulness labels and error explanations. This benchmark is for evaluation methods, testing whether a given method can detect challenging inconsistencies. Using this dataset, we first show that any one human annotation protocol is likely to miss inconsistencies, and we advocate for pursuing a range of methods when establishing ground truth for a summarization dataset. We finally test recent automatic metrics and find that none of them achieve more than 70% balanced accuracy on this task, demonstrating that it is a challenging benchmark for future work in faithfulness evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Agent-as-Judge for Factual Summarization of Long NarrativesYeonseok Jeong, Minsoo Kim, Seung-won Hwang, Byung-Hak KimEMNLP 2025 · 被引用 1 次
- Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction DetectionYuhao Chen, Yuanjie Lyu, Shuochen Liu, Chao Zhang 等EMNLP 2025
- Is the Top Still Spinning? Evaluating Subjectivity in Narrative UnderstandingMelanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang 等EMNLP 2025
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang 等ICLR 2024 · 被引用 176 次
相关 Paper
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 被引用 90 次
- Analyzing and Evaluating Faithfulness in Dialogue SummarizationBin Wang, Chen Zhang, Yan Zhang, Yiming Chen 等EMNLP 2022 · 被引用 15 次
- BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness MetricsLiang Ma, Shuyang Cao, Robert L. Logan IV, Di Lu 等ACL 2023
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai 等ACL 2024
- Questioning the Validity of Summarization Datasets and Improving Their Factual ConsistencyYanzhu Guo, Chloé Clavel, Moussa Kamal Eddine, Michalis VazirgiannisEMNLP 2022 · 被引用 5 次
