Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model
Qi Jia, Siyu Ren, Yizhu Liu, Kenny Q. Zhu
摘要
Despite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue. Previous work evaluates faithfulness either using models trained on the other tasks or in-domain synthetic data, or prompting a large model such as ChatGPT. This paper proposes to do zero-shot faithfulness evaluation simply with a moderately-sized foundation language model. We introduce a new metric FFLM, which is a combination of probability changes based on the intuition that prefixing a piece of text that is consistent with the output will increase the probability of predicting the output. Experiments show that FFLM performs competitively with or even outperforms ChatGPT on both inconsistency detection and faithfulness rating with 24x fewer parameters. FFLM also achieves improvements over other strong baselines. * The corresponding author. 1 We use the words "faithfulness", "consistency" and "(without) hallucination" interchangeably. Extrinsic hallucinations that are correct to the world knowledge are regarded as unfaithfulness in this work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software EngineeringRuiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan 等ISSTA 2025 · 被引用 27 次
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
- SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive SummarizationNayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu 等ACL 2025 · 被引用 7 次
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long ContextsYifei Yu, Qian-Wen Zhang, Lingfeng Qiao, Di Yin 等EMNLP 2025 · 被引用 2 次
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke 等ICLR 2025
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
相关 Paper
- Detecting and Mitigating Hallucinations in Multilingual SummarisationYifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti 等EMNLP 2023 · 被引用 12 次
- Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationSanjana Ramprasad, Byron C. WallaceNeurIPS 2025 · 被引用 13 次
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams 等EMNLP 2024 · 被引用 2 次
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai 等ACL 2024
- Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination TrendsSanjana Ramprasad, Elisa Ferracane, Zachary C. LiptonACL 2024
