Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation
Guocong Li, Qirui Hu, Ping Wang, Guofeng Zhang, Jian Wu, Hongxia Xu
摘要
Large Language Models enhanced with Retrieval Augmented Generation show strong potential in knowledge intensive tasks. However, they often encounter knowledge conflicts, where retrieved information contradicts the model's internal knowledge or exhibits internal inconsistencies. Existing methods force models into a binary choice between context and memory, leading to unreliable predictions. We argue that a more principled approach is to embrace contradictions as opportunities for deeper reasoning. To this end, we introduce Debate-of-Thoughts (DoT), a framework that transforms conflict resolution into an active deliberation process. DoT guides a single model through three phases: 1) hypothesis generation, which forms competing perspectives; 2) internal debate, where the model acts as both a proponent and a critic to stress test each view; and 3) adjudication, where the model acts as a judge to evaluate arguments based on evidence and logical consistency. We implement DoT via two complementary strategies: inference time prompt chaining and supervised fine tuning. Experiments across multiple conflict benchmarks show that DoT consistently outperforms state-of-the-art methods, while generating transparent debate transcripts that explain its decisions. By improving both accuracy and interpretability under knowledge conflicts, DoT establishes a more reliable paradigm for retrieval augmented generation systems. 1 * Corresponding authors. 1 Code are availabe at: https://github.com/cong03/ DoT . Contextual Conflict My internal data clearly indicates that ENIAC was completed in 1945 and is not the first commercial computer. User Query: Who is the current official men's marathon world record holder, and what is the time?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang 等EMNLP 2024 · 被引用 177 次
- Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and MethodTianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu 等SIGIR 2026
- Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented GenerationHua Ye, Siyuan Chen, Ziqi Zhong, Canran Xiao 等AAAI 2026 · 被引用 1 次
- Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented GenerationHaiyan Wu, Chenchen Wang, Chaoqun Sun, Chengxiong Lu 等WWW 2026
- Disentangling Reasoning Logic to Resolve Explicit Knowledge ConflictsXianda Zheng, Zijian Huang, Meng-Fen Chiang, Jiamou Liu 等ACL 2026 · 被引用 2 次
