MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
Justin Chih-Yao Chen, Archiki Prasad, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal
Abstract
Large language model (LLM) reasoning can be improved by scaling test-time compute with aggregation, i.e., generating multiple samples and aggregating over them. While improving performance, this strategy often reaches a saturation point beyond which additional compute provides no return. Refinement offers an alternative by using model-generated feedback to improve answer quality. However, refinement faces three key challenges: (1) Excessive refinement: Uniformly refining all instances can cause over-correction and reduce overall performance. (2) Inability to localize and address errors: LLMs struggle to identify and correct their own mistakes. (3) Insufficient refinement: Stopping refinement too soon could leave errors unaddressed. To tackle these issues, we propose MAGICORE, a framework for Multi-Agent Iteration for Coarse-to-fine Refinement. MAGICORE mitigates excessive refinement by categorizing problems as easy or hard, solving easy problems with coarsegrained aggregation, and solving the hard ones with fine-grained multi-agent refinement. To better localize errors, we incorporate external step-wise reward model scores, and to ensure sufficient refinement, we iteratively refine the solutions using a multi-agent setup. We evaluate MAGICORE on Llama-3-8B and GPT-3.5 and show its effectiveness across seven reasoning datasets. One iteration of MAGI-CORE beats Self-Consistency by 3.4%, Bestof-k by 3.2%, and Self-Refine by 4.0% even when these baselines use k = 120, and MAGI-CORE uses less than 50% of the compute. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69bef634-1346-4a1d-8e4c-9cc2964a9700Cited by top-tier papers5
- S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningRuotian Ma, Peisong Wang, Cheng Liu, Xingyan Liu et al.ACL 2025 · 13 citations
- PRInTS: Reward Modeling for Long-Horizon Information SeekingJaewoo Lee, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan et al.ACL 2026 · 3 citations
- The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden RepresentationsYubo Zhu, Dongrui Liu, Zecheng Lin, Wei Tong et al.EMNLP 2025 · 1 citation
- Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy VerifierYegor Denisov-Blanch, Joshua Kazdan, Jessica Chudnovsky, Rylan Schaeffer et al.ICML 2026
- OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical ModelingMaxime Bouscary, Saurabh AminAAAI 2026
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsDeepak Nathani, David Wang, Liangming Pan, William Yang WangEMNLP 2023 · 6 citations
- Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi et al.ACL 2025 · 13 citations
- MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQLArian Askari, Christian Pölitz, Xinye TangAAAI 2025 · 44 citations
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use MixtureYongchao Chen, Jiefeng Chen, Rui Meng, Ji Yin et al.ICLR 2026 · 13 citations
- D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool UseBowen Xu, Shaoyu Wu, Hao Jiang, Kai Liu et al.ICML 2026
