Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization
Rongxin Zhu, Jianzhong Qi, Jey Han Lau
摘要
A series of datasets and models have been proposed for summaries generated for wellformatted documents such as news articles. Dialogue summaries, however, have been under explored. In this paper, we present the first dataset with fine-grained factual error annotations named DIASUMFACT. We define finegrained factual error detection as a sentencelevel multi-label classification problem, and we evaluate two state-of-the-art (SOTA) models on our dataset. Both models yield sub-optimal results, with a macro-averaged F1 score of around 0.25 over 6 error classes. We further propose an unsupervised model ENDERANKER via candidate ranking using pretrained encoder-decoder models. Our model performs on par with the SOTA models while requiring fewer resources. These observations confirm the challenges in detecting factual errors from dialogue summaries, which call for further studies, for which our dataset and results offer a solid foundation. 1 Lilly: Wanna go out tonight? Marshall: can't :( money's low Lilly: my treat :) Marshall: I wouldn't let a woman pay for me.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement LearningShuzheng Si, Haozhe Zhao, Cheng Gao, Yuzhuo Bai 等AAAI 2026 · 被引用 4 次
- HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought ReasoningShayan Ali Akbar, Md Mosharaf Hossain, Tess Wood, Si-Chi Chin 等EMNLP 2024 · 被引用 4 次
- Dialogue Summarization with Mixture of Experts based on Large Language ModelsYuanhe Tian, Fei Xia, Yan SongACL 2024
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking StrategiesYuxuan Ye, Raúl Santos-Rodríguez, Edwin SimpsonACL 2026
- Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination TrendsSanjana Ramprasad, Elisa Ferracane, Zachary C. LiptonACL 2024
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
相关 Paper
- Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation FrameworkMingqi Gao, Xiaojun Wan, Jia Su, Zhefeng Wang 等ACL 2023 · 被引用 4 次
- Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual ErrorsAlex Chandler, Devesh Surve, Hui SuEMNLP 2024 · 被引用 2 次
- DialFact: A Benchmark for Fact-Checking in DialoguePrakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming XiongACL 2022
- Analyzing and Evaluating Faithfulness in Dialogue SummarizationBin Wang, Chen Zhang, Yan Zhang, Yiming Chen 等EMNLP 2022 · 被引用 15 次
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban 等ACL 2023 · 被引用 38 次
