Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends
Sanjana Ramprasad, Elisa Ferracane, Zachary C. Lipton
Abstract
Recent advancements in large language models (LLMs) have considerably advanced the capabilities of summarization systems. However, they continue to face concerns about hallucination. While prior work has evaluated LLMs extensively in news domains, most evaluation of dialogue summarization has focused on BART-based models, leaving a gap in our understanding of their faithfulness. Our work benchmarks the faithfulness of LLMs for dialogue summarization, using human annotations and focusing on identifying and categorizing span-level inconsistencies. Specifically, we focus on two prominent LLMs: GPT-4 and Alpaca-13B. Our evaluation reveals subtleties as to what constitutes a hallucination: LLMs often generate plausible inferences, supported by circumstantial evidence in the conversation, that lack direct evidence, a pattern that is less prevalent in older models. We propose a refined taxonomy of errors, coining the category of "Circumstantial Inference" to bucket these LLM behaviors. Using our taxonomy, we compare the behavioral differences between LLMs and older fine-tuned models. Additionally, we systematically assess the efficacy of automatic error detection methods on LLM summaries and find that they struggle to detect these nuanced errors. To address this, we introduce two prompt-based approaches for fine-grained error detection that outperform existing metrics, particularly for identifying "Circumstantial Inference." 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f87e55ca-a4cf-4c46-9b8b-89626bdc7daeCited by top-tier papers8
- Summaries, Highlights, and Action Items: Design, Implementation and Evaluation of an LLM-powered Meeting Recap SystemSumit Asthana, Sagih Hilleli, Pengcheng He, Aaron HalfakerCSCW 2025 · 24 citations
- Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationSanjana Ramprasad, Byron C. WallaceNeurIPS 2025 · 13 citations
- Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat LogsRupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger et al.ACL 2025 · 3 citations
- CoHear: Conversation Enhancement via Multi-earphone CollaborationLixing He, Yunqi Guo, Zhenyu Yan, Guoliang XingUbiComp 2026 · 1 citation
- DioR: Adaptive Cognitive Detection and Contextual Retrieval Optimization for Dynamic Retrieval-Augmented GenerationHanghui Guo, Jia Zhu, Shimin Di, Weijie Shi et al.ACL 2025
Builds on9
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 121 citations
- What Have We Achieved on Text Summarization?Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao et al.EMNLP 2020 · 72 citations
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban et al.ACL 2023 · 38 citations
- Analyzing and Evaluating Faithfulness in Dialogue SummarizationBin Wang, Chen Zhang, Yan Zhang, Yiming Chen et al.EMNLP 2022 · 15 citations
Related papers
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
- Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention MapsYung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna et al.EMNLP 2024 · 18 citations
- HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought ReasoningShayan Ali Akbar, Md Mosharaf Hossain, Tess Wood, Si-Chi Chin et al.EMNLP 2024 · 4 citations
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri et al.EMNLP 2023 · 29 citations
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelQi Jia, Siyu Ren, Yizhu Liu, Kenny Q. ZhuEMNLP 2023 · 4 citations
