Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
Tianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu, Qika Lin, Jun Liu
Abstract
Large language models (LLMs) have achieved remarkable success across a wide range of applications especially when augmented by external knowledge through retrieval-augmented generation (RAG). Despite their widespread adoption, recent studies have shown that LLMs often struggle to perform faithful reasoning when conflicting knowledge is retrieved. However, existing work primarily focuses on conflicts between external knowledge and the parametric knowledge of LLMs, leaving conflicts across external knowledge largely unexplored. Meanwhile, modern RAG systems increasingly emphasize the integration of unstructured text and (semi-)structured data like knowledge graphs (KGs) to improve knowledge completeness and reasoning faithfulness. To address this gap, we introduce ConflictQA, a novel benchmark that systematically instantiates conflicts between textual evidence and KG evidence. Extensive evaluations across representative LLMs reveal that, facing such crosssource conflicts, LLMs often fail to identify reliable evidence for correct reasoning. Instead, LLMs become more sensitive to prompting choices and tend to rely exclusively on either KG or textual evidence, resulting in incorrect responses. Based on these findings, we further propose XoT, a two-stage explanation-based thinking framework tailored for reasoning over heterogeneous conflicting evidence, and verify its effectiveness with extensive experiments.
• Computing methodologies → Reasoning about belief and knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 630d79df-e38c-4b44-beb3-8b8251b334b2Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous SourcesXingxuan Li, Ruochen Zhao, Yew Ken Chia, Bosheng Ding et al.ICLR 2024 · 165 citations
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 134 citations
- Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsFei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen et al.ACL 2025 · 50 citations
Related papers
- TruthfulRAG: Resolving Factual-level Conflicts in Retrieval-Augmented Generation with Knowledge GraphsShuyi Liu, Yu-Ming Shang, Xi ZhangAAAI 2026 · 2 citations
- Empowering GraphRAG with Knowledge Filtering and IntegrationKai Guo, Harry Shomer, Shenglai Zeng, Haoyu Han et al.EMNLP 2025 · 2 citations
- FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented GenerationQinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang et al.ACL 2025 · 18 citations
- Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesYizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li et al.ACL 2026
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsYifan Jia, Yuntao Du, Kailin Jiang, Yuyang Liang et al.AAAI 2026
