Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization
Rongxin Zhu, Jianzhong Qi, Jey Han Lau
Abstract
A series of datasets and models have been proposed for summaries generated for wellformatted documents such as news articles. Dialogue summaries, however, have been under explored. In this paper, we present the first dataset with fine-grained factual error annotations named DIASUMFACT. We define finegrained factual error detection as a sentencelevel multi-label classification problem, and we evaluate two state-of-the-art (SOTA) models on our dataset. Both models yield sub-optimal results, with a macro-averaged F1 score of around 0.25 over 6 error classes. We further propose an unsupervised model ENDERANKER via candidate ranking using pretrained encoder-decoder models. Our model performs on par with the SOTA models while requiring fewer resources. These observations confirm the challenges in detecting factual errors from dialogue summaries, which call for further studies, for which our dataset and results offer a solid foundation. 1 Lilly: Wanna go out tonight? Marshall: can't :( money's low Lilly: my treat :) Marshall: I wouldn't let a woman pay for me.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53e7304c-891a-4108-aab5-5d0bd852153cCited by top-tier papers5
- Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement LearningShuzheng Si, Haozhe Zhao, Cheng Gao, Yuzhuo Bai et al.AAAI 2026 · 4 citations
- HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought ReasoningShayan Ali Akbar, Md Mosharaf Hossain, Tess Wood, Si-Chi Chin et al.EMNLP 2024 · 4 citations
- Dialogue Summarization with Mixture of Experts based on Large Language ModelsYuanhe Tian, Fei Xia, Yan SongACL 2024
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking StrategiesYuxuan Ye, Raúl Santos-Rodríguez, Edwin SimpsonACL 2026
- Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination TrendsSanjana Ramprasad, Elisa Ferracane, Zachary C. LiptonACL 2024
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
Related papers
- Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation FrameworkMingqi Gao, Xiaojun Wan, Jia Su, Zhefeng Wang et al.ACL 2023 · 4 citations
- Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual ErrorsAlex Chandler, Devesh Surve, Hui SuEMNLP 2024 · 2 citations
- DialFact: A Benchmark for Fact-Checking in DialoguePrakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming XiongACL 2022
- Analyzing and Evaluating Faithfulness in Dialogue SummarizationBin Wang, Chen Zhang, Yan Zhang, Yiming Chen et al.EMNLP 2022 · 15 citations
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban et al.ACL 2023 · 38 citations
