COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated Texts
Aviya Maimon, Reut Tsarfaty
摘要
Coherence is a linguistic term that refers to the relations between small textual units (sentences, propositions), which make the text logically consistent and meaningful to the reader. With the advances of generative foundational models in NLP, there is a pressing need to automatically assess the human-perceived coherence of automatically generated texts. Up until now, little work has been done on explicitly assessing the coherence of generated texts and analyzing the factors contributing to (in)coherence. Previous work on the topic used other tasks, e.g., sentence reordering, as proxies of coherence, rather than approaching coherence detection heads on. In this paper, we introduce CoheSentia, a novel benchmark of human-perceived coherence of automatically generated texts. Our annotation protocol reflects two perspectives; one is global, assigning a single coherence score, and the other is incremental, scoring sentence by sentence. The incremental method produces an (in)coherence score for each text fragment and also pinpoints reasons for incoherence at that point. Our benchmark contains 500 automatically-generated and human-annotated paragraphs, each annotated in both methods, by multiple raters. Our analysis shows that the inter-annotator agreement in the incremental mode is higher than in the holistic alternative, and our experiments show that standard LMs fine-tuned for coherence detection show varied performance on the different factors contributing to (in)coherence. All in all, these models yield unsatisfactory performance, emphasizing the need for developing more reliable methods for coherence assessment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and RewritingXuanming Zhang, Anthony Diaz, Zixun Chen, Qingyang Wu 等EMNLP 2024 · 被引用 2 次
- What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story EvaluationDingyi Yang, Qin JinACL 2025 · 被引用 2 次
- From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM ReasoningSeungdong Yoa, Sanghyu Yoon, Suhee Yoon, Dongmin Kim 等ICLR 2026 · 被引用 1 次
- Joint Modeling of Entities and Discourse Relations for Coherence AssessmentWei Liu, Michael StrubeEMNLP 2025
它引用的顶会 Paper1
相关 Paper
- BBScore: A Brownian Bridge Based Metric for Assessing Text CoherenceZhecheng Sheng, Tianhao Zhang, Chen Jiang, Dongyeop KangAAAI 2024 · 被引用 8 次
- Hierarchical Coherence Modeling for Document Quality AssessmentDongliang Liao, Jin Xu, Gongfu Li, Yiru WangAAAI 2021 · 被引用 19 次
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
- Coherence boosting: When your pretrained language model is not paying enough attentionNikolay Malkin, Zhen Wang, Nebojsa JojicACL 2022 · 被引用 45 次
- Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language ModelsDan Iter, Kelvin Guu, Larry Lansing, Dan JurafskyACL 2020 · 被引用 72 次
