Analyzing and Evaluating Faithfulness in Dialogue Summarization
Bin Wang, Chen Zhang, Yan Zhang, Yiming Chen, Haizhou Li
摘要
Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text summarization. However, there is a lack of systematic study on dialogue summarization systems. In this work, we first perform the fine-grained human analysis on the faithfulness of dialogue summaries and observe that over 35% of generated summaries are faithfully inconsistent respective the source dialogues. Furthermore, we present a new model-level faithfulness evaluation method. It examines generation models with multi-choice questions created by rule-based transformations. Experimental results show that our evaluation schema is a strong proxy for the factual correctness of summarization models. The humanannotated faithfulness samples and the evaluation toolkit are released to facilitate future research toward faithful dialogue summarization. Code available: https://github. com/BinWang28/FacEval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationSanjana Ramprasad, Byron C. WallaceNeurIPS 2025 · 被引用 13 次
- Instructive Dialogue Summarization with Query AggregationsBin Wang, Zhengyuan Liu, Nancy F. ChenEMNLP 2023 · 被引用 5 次
- Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation FrameworkMingqi Gao, Xiaojun Wan, Jia Su, Zhefeng Wang 等ACL 2023 · 被引用 4 次
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?Ashutosh Bajpai, Tanmoy ChakrabortyEMNLP 2025
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang 等ACL 2025
它引用的顶会 Paper15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang 等ACL 2020 · 被引用 410 次
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 被引用 329 次
相关 Paper
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams 等EMNLP 2024 · 被引用 2 次
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 被引用 90 次
- X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive SummarizationSubhajit Chaudhury, Sarathkrishna Swaminathan, R. Chulaka Gunasekara, Maxwell Crouse 等EMNLP 2022 · 被引用 13 次
- Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive SummarizationFaisal Ladhak, Esin Durmus, He He, Claire Cardie 等ACL 2022 · 被引用 74 次
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
