Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation
Guocong Li, Qirui Hu, Ping Wang, Guofeng Zhang, Jian Wu, Hongxia Xu
Abstract
Large Language Models enhanced with Retrieval Augmented Generation show strong potential in knowledge intensive tasks. However, they often encounter knowledge conflicts, where retrieved information contradicts the model's internal knowledge or exhibits internal inconsistencies. Existing methods force models into a binary choice between context and memory, leading to unreliable predictions. We argue that a more principled approach is to embrace contradictions as opportunities for deeper reasoning. To this end, we introduce Debate-of-Thoughts (DoT), a framework that transforms conflict resolution into an active deliberation process. DoT guides a single model through three phases: 1) hypothesis generation, which forms competing perspectives; 2) internal debate, where the model acts as both a proponent and a critic to stress test each view; and 3) adjudication, where the model acts as a judge to evaluate arguments based on evidence and logical consistency. We implement DoT via two complementary strategies: inference time prompt chaining and supervised fine tuning. Experiments across multiple conflict benchmarks show that DoT consistently outperforms state-of-the-art methods, while generating transparent debate transcripts that explain its decisions. By improving both accuracy and interpretability under knowledge conflicts, DoT establishes a more reliable paradigm for retrieval augmented generation systems. 1 * Corresponding authors. 1 Code are availabe at: https://github.com/cong03/ DoT . Contextual Conflict My internal data clearly indicates that ENIAC was completed in 1945 and is not the first commercial computer. User Query: Who is the current official men's marathon world record holder, and what is the time?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
- Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and MethodTianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu et al.SIGIR 2026
- Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented GenerationHua Ye, Siyuan Chen, Ziqi Zhong, Canran Xiao et al.AAAI 2026 · 1 citation
- Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented GenerationHaiyan Wu, Chenchen Wang, Chaoqun Sun, Chengxiong Lu et al.WWW 2026
- Disentangling Reasoning Logic to Resolve Explicit Knowledge ConflictsXianda Zheng, Zijian Huang, Meng-Fen Chiang, Jiamou Liu et al.ACL 2026 · 2 citations
