Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
Tianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu, Qika Lin, Jun Liu
摘要
Large language models (LLMs) have achieved remarkable success across a wide range of applications especially when augmented by external knowledge through retrieval-augmented generation (RAG). Despite their widespread adoption, recent studies have shown that LLMs often struggle to perform faithful reasoning when conflicting knowledge is retrieved. However, existing work primarily focuses on conflicts between external knowledge and the parametric knowledge of LLMs, leaving conflicts across external knowledge largely unexplored. Meanwhile, modern RAG systems increasingly emphasize the integration of unstructured text and (semi-)structured data like knowledge graphs (KGs) to improve knowledge completeness and reasoning faithfulness. To address this gap, we introduce ConflictQA, a novel benchmark that systematically instantiates conflicts between textual evidence and KG evidence. Extensive evaluations across representative LLMs reveal that, facing such crosssource conflicts, LLMs often fail to identify reliable evidence for correct reasoning. Instead, LLMs become more sensitive to prompting choices and tend to rely exclusively on either KG or textual evidence, resulting in incorrect responses. Based on these findings, we further propose XoT, a two-stage explanation-based thinking framework tailored for reasoning over heterogeneous conflicting evidence, and verify its effectiveness with extensive experiments.
• Computing methodologies → Reasoning about belief and knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
- Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous SourcesXingxuan Li, Ruochen Zhao, Yew Ken Chia, Bosheng Ding 等ICLR 2024 · 被引用 165 次
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 被引用 134 次
- Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsFei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen 等ACL 2025 · 被引用 50 次
相关 Paper
- TruthfulRAG: Resolving Factual-level Conflicts in Retrieval-Augmented Generation with Knowledge GraphsShuyi Liu, Yu-Ming Shang, Xi ZhangAAAI 2026 · 被引用 2 次
- Empowering GraphRAG with Knowledge Filtering and IntegrationKai Guo, Harry Shomer, Shenglai Zeng, Haoyu Han 等EMNLP 2025 · 被引用 2 次
- FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented GenerationQinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang 等ACL 2025 · 被引用 18 次
- Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesYizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li 等ACL 2026
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsYifan Jia, Yuntao Du, Kailin Jiang, Yuyang Liang 等AAAI 2026
