Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction
Payel Santra, Lavisha Sharma, Madhusudan Ghosh, Partha Basuchowdhuri
Abstract
The rapid spread of misinformation on social media highlights the need for robust, automated fact correction frameworks. However, existing works rely on supervised learning from manually annotated claim-evidence pairs, which are scarce and prone to biases, limiting their generalization across domains. Moreover, these methods overlook semantic faithfulness in their correction process. To address these challenges, we propose Mask-to-Correct (MC), a training-free, inference-only Retrieval Augmented Generation (RAG) based framework that leverages diversity-aware masking to identify erroneous spans of claims and evaluate the faithfulness of corrections using retrieved evidence. However, the effectiveness of RAG heavily depends on the choice of retriever, which may vary across queries. To mitigate this, we further introduce MC, an ensemble-based framework that combines corrections across multiple rankers to reduce retrieval bias and improve robustness. Extensive experiments on the benchmark datasets demonstrate that our proposed frameworks consistently outperform all baselines, achieving up to 14% improvement in SARI scores, without using gold evidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4f3518f-7ce8-4804-b475-993d0f26eb82Builds on26
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
Related papers
- SeaRAG: Reducing Hallucination in Retrieval-Augmented Generation via Statement-Entity Adaptive RankingXiaosong Yuan, Xiaofeng Zhang, Di Zhao, Yijia Zhang et al.WWW 2026
- Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-CheckingShuzhi Gong, Richard O. Sinnott, Jianzhong Qi, Cécile Paris et al.SIGIR 2026 · 2 citations
- Removal of Hallucination on Hallucination: Debate-Augmented RAGWentao Hu, Wengyu Zhang, Yiyang Jiang, Chen Jason Zhang et al.ACL 2025
- Counterfactual Debiasing for Fact VerificationWeizhi Xu, Qiang Liu, Shu Wu, Liang WangACL 2023 · 26 citations
- Improving Factual Error Correction by Learning to Inject Factual ErrorsXingwei He, Qianru Zhang, A-Long Jin, Jun Ma et al.AAAI 2024 · 5 citations
