REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
Chuyi Kong, Wei Gao, Jing Ma, Hongzhan Lin, Yuxi Sun
Abstract
The prevalence of fake news on social media calls for automated fact-checking systems that deliver not only accurate verdicts but also faithful explanations. However, existing large language model (LLM)-based methods often overlook deceptive misinformation styles in generated explanations, producing unfaithful rationales that may mislead human judgment. They also rely heavily on external knowledge sources, which can introduce hallucinations and incur substantial latency, undermining both reliability and responsiveness in realtime settings. To address these limitations, we propose REason-guided Fact-checking with Latent EXplanations (REFLEX), a selfrefining framework that explicitly controls reasoning style by anchoring explanations to the predicted verdict. REFLEX leverages selfdisagreement veracity signals between a backbone model and its fine-tuned variant to construct steering vectors, thereby naturally disentangling factual content from stylistic cues. Experiments on a real-world benchmark show that REFLEX achieves state-of-the-art performance under LLaMA-series models using only 465 self-refined samples. Owing to its transferability, REFLEX also yields gains of up to 7.54 Macro-F1 points on in-the-wild data. Further analysis shows that our method effectively mitigates faithful hallucination, leading to both more reliable explanations and more accurate verdicts than prior explainable fact-checking approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b39d58d2-775a-47a4-b004-c8d3f9b5ac74Builds on23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- ReFL: Reflective Feedback Learning for Hallucination Detection of Large Language ModelsCunhang Fan, Jun Zhang, Xue Zhang, Shuai Zhang et al.ACL 2026
- A Fact-Checking Framework with Denoising Evidence Retrieval and LLM-Based Debate VerificationJun Yang, Yuhan Bai, Dandan Song, Zhijing Wu et al.WWW 2026
- Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionChaowei Zhang, Zongling Feng, Zewei Zhang, Jipeng Qiang et al.AAAI 2025 · 13 citations
- Improving Retrieval Augmented Language Model with Self-ReasoningYuan Xia, Jingbo Zhou, Zhenhui Shi, Jun Chen et al.AAAI 2025 · 42 citations
- Explainable Fake News Detection with Large Language Model via Defense Among Competing WisdomBo Wang, Jing Ma, Hongzhan Lin, Zhiwei Yang et al.WWW 2024 · 104 citations
