Zero-shot Faithful Factual Error Correction
Kung-Hsiang Huang, Hou Pong Chan, Heng Ji
Abstract
Faithfully correcting factual errors is critical for maintaining the integrity of textual knowledge bases and preventing hallucinations in sequence-to-sequence models. Drawing on humans' ability to identify and correct factual errors, we present a zero-shot framework that formulates questions about input claims, looks for correct answers in the given evidence, and assesses the faithfulness of each correction based on its consistency with the evidence. Our zero-shot framework outperforms fully-supervised approaches, as demonstrated by experiments on the FEVER and SciFact datasets, where our outputs are shown to be more faithful. More importantly, the decomposability nature of our framework inherently provides interpretability. Additionally, to reveal the most suitable metrics for evaluating factual error corrections, we analyze the correlation between commonly used metrics with human judgments in terms of three different dimensions regarding intelligibility and faithfulness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 966336ec-47d1-47f9-ba3a-2d202d5f3ad8Cited by top-tier papers7
- The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language ModelsJingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu et al.EMNLP 2023 · 20 citations
- NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyYi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow et al.EMNLP 2023 · 17 citations
- SafeWorld: Geo-Diverse Safety AlignmentDa Yin, Haoyi Qiu, Kung-Hsiang Huang, Kai-Wei Chang et al.NeurIPS 2024 · 14 citations
- Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response ForecastingChenkai Sun, Jinning Li, Yi Ren Fung, Hou Pong Chan et al.EMNLP 2023 · 12 citations
- Gender Biases in Automatic Evaluation Metrics for Image CaptioningHaoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz et al.EMNLP 2023 · 6 citations
Builds on12
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationShuyang Cao, Lu WangEMNLP 2021 · 130 citations
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 67 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
Related papers
- Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text GenerationTathagata Raha, Clément Christophe, Nada Saadi, Hamza Javed et al.ACL 2026
- LM vs LM: Detecting Factual Errors via Cross ExaminationRoi Cohen, May Hamri, Mor Geva, Amir GlobersonEMNLP 2023 · 41 citations
- HalluClean: A Unified Framework to Combat Hallucinations in LLMsYaxin Zhao, Yu ZhangAAAI 2026
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelQi Jia, Siyu Ren, Yizhu Liu, Kenny Q. ZhuEMNLP 2023 · 4 citations
