SNaC: Coherence Error Detection for Narrative Summarization
Tanya Goyal, Junyi Jessy Li, Greg Durrett
Abstract
Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks. A long summary that appropriately covers the facets of that text must also present a coherent narrative, but current automatic and human evaluation methods fail to identify gaps in coherence. In this work, we introduce SNAC, a narrative coherence evaluation framework for fine-grained annotations of long summaries. We develop a taxonomy of coherence errors in generated narrative summaries and collect spanlevel annotations for 6.6k sentences across 150 book and movie summaries. Our work provides the first characterization of coherence errors generated by state-of-the-art summarization models and a protocol for eliciting coherence judgments from crowdworkers. Furthermore, we show that the collected annotations allow us to benchmark past work in coherence modeling and train a strong classifier for automatically localizing coherence errors in generated summaries. Finally, our SNAC framework can support future work in long document summarization and coherence evaluation, including improved summarization modeling and posthoc summary correction. 1 Corresponding excerpt from the human-written summary properly contextualizes the new character. Recently, Wu et al. (2021) proposed a strong book summarization model but showed that although generated summaries covered important information from the books, they read like a list of events stapled together without any coherent narrative structure (see Figure 1 ). We found similar Context: John Fenwick, an aspiring artist, accepts a loan from Mr. Morrison to move to London to pursue his art career. In London, he becomes infatuated with Madame de Pastourelles, a beautiful and intelligent artist.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b35753b7-f19f-4303-b226-8da2db7b96a9Cited by top-tier papers8
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
- Leveraging Locality in Abstractive Text SummarizationYixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb et al.EMNLP 2022 · 19 citations
- Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive SummarizationShiyue Zhang, David Wan, Mohit BansalACL 2023 · 15 citations
- Concise Answers to Complex Questions: Summarization of Long-form AnswersAbhilash Potluri, Fangyuan Xu, Eunsol ChoiACL 2023 · 4 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang et al.ACL 2022 · 62 citations
Related papers
- NexusSum: Hierarchical LLM Agents for Long-Form Narrative SummarizationHyuntak Kim, Byung-Hak KimACL 2025
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams et al.EMNLP 2024 · 2 citations
- BOOKCOREF: Coreference Resolution at Book ScaleGiuliano Martinelli, Tommaso Bonomo, Pere-Lluís Huguet Cabot, Roberto NavigliACL 2025
- Agent-as-Judge for Factual Summarization of Long NarrativesYeonseok Jeong, Minsoo Kim, Seung-won Hwang, Byung-Hak KimEMNLP 2025 · 1 citation
- LiteraryQA: Towards Effective Evaluation of Long-document Narrative QATommaso Bonomo, Luca Gioffré, Roberto NavigliEMNLP 2025 · 1 citation
