LM2: A Simple Society of Language Models Solves Complex Reasoning
Gurusha Juneja, Subhabrata Dutta, Tanmoy Chakraborty
Abstract
Despite demonstrating emergent reasoning abilities, Large Language Models (LLMS) often lose track of complex, multi-step reasoning. Existing studies show that providing guidance via decomposing the original question into multiple subproblems elicits more robustness in LLM reasoning -a decomposer generates the subproblems, and a solver solves each of these subproblems. However, these techniques fail to accommodate coordination between the decomposer and the solver modules (either in a single model or different specialized ones) -the decomposer does not keep track of the ability of the solver to follow the decomposed reasoning. In this paper, we propose LM 2 to address these challenges. LM 2 modularizes the decomposition, solution, and verification into three different language models. The decomposer module identifies the key concepts necessary to solve the problem and generates step-by-step subquestions according to the reasoning requirement. The solver model generates the solution to the subproblems that are then checked by the verifier module; depending upon the feedback from the verifier, the reasoning context is constructed using the subproblems and the solutions. These models are trained to coordinate using policy learning. Exhaustive experimentation suggests the superiority of LM 2 over existing methods on in-and out-domain reasoning problems, outperforming the best baselines by 8.1% on MATH, 7.71% on JEEBench, and 9.7% on MedQA problems (code available at https://github.com/ LCS2-IIITD/Language_Model_Multiplex ). : How many distinct, non-equilateral triangles with a perimeter of 60 units have integer side lengths , , and such that , , is an arithmetic sequence? Solver LM Decomposer LM Verifier LM SQ: What is a, b, c in terms of common difference d? SA:Since a, b, and c form an arithmetic sequence, we can express c in terms of a as c = a + d, where d is the common difference Let be the common difference, so and We can assume that is positive In particular, can't be 0, because the triangle is not equilateral Then the perimeter of the triangle is , so Hence, the sides of the triangle are , 20, and These sides must satisfy the triangle inequality, which gives us Solving for , we find , or Therefore, the possible values of are 1, 2, , 9, which gives us possible triangles
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Emergent Coordination in Multi-Agent Language ModelsChristoph RiedlICLR 2026 · 25 citations
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
- PaT: Planning-after-Trial for Efficient Test-Time Code GenerationYoungsik Yoon, Sungjae Lee, Seockbean Song, Siwei Wang et al.ACL 2026
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu et al.ICLR 2023 · 94 citations
Related papers
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo et al.EMNLP 2023 · 2 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-TrainingLeonardo Ranaldi, Federico Ranaldi, Giulia PucciACL 2025 · 9 citations
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
