COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated Texts
Aviya Maimon, Reut Tsarfaty
Abstract
Coherence is a linguistic term that refers to the relations between small textual units (sentences, propositions), which make the text logically consistent and meaningful to the reader. With the advances of generative foundational models in NLP, there is a pressing need to automatically assess the human-perceived coherence of automatically generated texts. Up until now, little work has been done on explicitly assessing the coherence of generated texts and analyzing the factors contributing to (in)coherence. Previous work on the topic used other tasks, e.g., sentence reordering, as proxies of coherence, rather than approaching coherence detection heads on. In this paper, we introduce CoheSentia, a novel benchmark of human-perceived coherence of automatically generated texts. Our annotation protocol reflects two perspectives; one is global, assigning a single coherence score, and the other is incremental, scoring sentence by sentence. The incremental method produces an (in)coherence score for each text fragment and also pinpoints reasons for incoherence at that point. Our benchmark contains 500 automatically-generated and human-annotated paragraphs, each annotated in both methods, by multiple raters. Our analysis shows that the inter-annotator agreement in the incremental mode is higher than in the holistic alternative, and our experiments show that standard LMs fine-tuned for coherence detection show varied performance on the different factors contributing to (in)coherence. All in all, these models yield unsatisfactory performance, emphasizing the need for developing more reliable methods for coherence assessment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff4594e2-9ccf-4737-8f24-972b69c012e7Cited by top-tier papers4
- DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and RewritingXuanming Zhang, Anthony Diaz, Zixun Chen, Qingyang Wu et al.EMNLP 2024 · 2 citations
- What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story EvaluationDingyi Yang, Qin JinACL 2025 · 2 citations
- From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM ReasoningSeungdong Yoa, Sanghyu Yoon, Suhee Yoon, Dongmin Kim et al.ICLR 2026 · 1 citation
- Joint Modeling of Entities and Discourse Relations for Coherence AssessmentWei Liu, Michael StrubeEMNLP 2025
Builds on1
Related papers
- BBScore: A Brownian Bridge Based Metric for Assessing Text CoherenceZhecheng Sheng, Tianhao Zhang, Chen Jiang, Dongyeop KangAAAI 2024 · 8 citations
- Hierarchical Coherence Modeling for Document Quality AssessmentDongliang Liao, Jin Xu, Gongfu Li, Yiru WangAAAI 2021 · 19 citations
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
- Coherence boosting: When your pretrained language model is not paying enough attentionNikolay Malkin, Zhen Wang, Nebojsa JojicACL 2022 · 45 citations
- Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language ModelsDan Iter, Kelvin Guu, Larry Lansing, Dan JurafskyACL 2020 · 72 citations
