ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
Justin Chih-Yao Chen, Swarnadeep Saha, Mohit Bansal
Abstract
Large Language Models (LLMs) still struggle with natural language reasoning tasks. Motivated by the society of minds (Minsky, 1988) , we propose RECONCILE, a multi-model multiagent framework designed as a round table conference among diverse LLM agents. RECON-CILE enhances collaborative reasoning between LLM agents via multiple rounds of discussion, learning to convince other agents to improve their answers, and employing a confidenceweighted voting mechanism that leads to a better consensus. In each round, RECONCILE initiates discussion between agents via a 'discussion prompt' that consists of (a) grouped answers and explanations generated by each agent in the previous round, (b) their confidence scores, and (c) demonstrations of answerrectifying human explanations, used for convincing other agents. Experiments on seven benchmarks demonstrate that RECONCILE significantly improves LLMs' reasoning -both individually and as a team -surpassing prior single-agent and multi-agent baselines by up to 11.4% and even outperforming GPT-4 on three datasets. RECONCILE also flexibly incorporates different combinations of agents, including API-based, open-source, and domainspecific models, leading to an 8% improvement on MATH. Finally, we analyze the individual components of RECONCILE, demonstrating that the diversity originating from different models is critical to its superior performance. 1 Self-Refine MAD+Judge Multi-Agent Debate (MAD) ReConcile (Group-Discuss-and-Convince) Yes, with 95% confidence No, with 50% confidence No, with 40% confidence yes no no yes no no yes no no Question (Q): Is an ammonia fighting cleaner good for pet owners? Human Explanation (Exp): Ammonia is a component in pet urine. It has an unpleasant odor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe2f38b1-db7e-43ae-9eda-6bc8296d7f51Cited by top-tier papers44
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Chain of Agents: Large Language Models Collaborating on Long-Context TasksYusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister et al.NeurIPS 2024 · 297 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- Multi-Agent Design: Optimizing Agents with Better Prompts and TopologiesHan Zhou, Xingchen Wan, Ruoxi Sun, Hamid Palangi et al.ICLR 2026 · 127 citations
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?Hyeong Kyu Choi, Xiaojin Zhu, Sharon LiNeurIPS 2025 · 93 citations
Builds on26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang et al.ICLR 2026
- Multi-Agent Debate with Memory MaskingHongduan Tian, Xiao Feng, Ziyuan Zhao, Xiangyu Zhu et al.ICLR 2026 · 7 citations
- Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent DebateYexiang Liu, Jie Cao, Zekun Li, Ran He et al.ICLR 2025
- Collaborative Reasoner: Self-Improving Social Agents with Synthetic ConversationsAnsong Ni, Ruta Desai, Yang Li, Xinjie Lei et al.NeurIPS 2025 · 7 citations
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
