Debating with More Persuasive LLMs Leads to More Truthful Answers
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, Ethan Perez
摘要
Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information but are otherwise as capable. The method we evaluate is debate, where two LLM experts each argue for a different answer, and a non-expert selects the answer. On the QuALITY comprehension task, we find that debate consistently helps both non-expert models and humans answer questions, achieving 76% and 88% accuracy respectively (naive baselines obtain 48% and 60%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper72
- Multi-LLM Debate: Framework, Principals, and InterventionsAndrew Estornell, Yang LiuNeurIPS 2024 · 被引用 131 次
- Multi-Agent Design: Optimizing Agents with Better Prompts and TopologiesHan Zhou, Xingchen Wan, Ruoxi Sun, Hamid Palangi 等ICLR 2026 · 被引用 127 次
- On scalable oversight with weak LLMs judging strong LLMsZachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen 等NeurIPS 2024 · 被引用 116 次
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange 等ICLR 2026 · 被引用 101 次
- GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual ReasoningJusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen 等NeurIPS 2025 · 被引用 75 次
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
相关 Paper
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett 等ICML 2024 · 被引用 82 次
- Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less UsefulChenyan Jia, Apoorva Gondimalla, Angie Zhang, David Joseph Mullings 等CHI 2026 · 被引用 1 次
- Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacksVirgile Rennard, Christos Xypolopoulos, Michalis VazirgiannisACL 2025 · 被引用 8 次
- Grounded in Reality: Learning and Deploying Proactive LLM from Offline LogsFei Wei, Daoyuan Chen, Ce Wang, Yilun Huang 等ICML 2026 · 被引用 2 次
