ThinkSLM: Towards Reasoning in Small Language Models
Gaurav Srivastava, Shuxiang Cao, Xuan Wang
Abstract
Reasoning has long been viewed as an emergent property of large language models (LLMs). However, recent studies challenge this assumption, showing that small language models (SLMs) can also achieve competitive reasoning performance. This paper introduces THINKSLM, the first extensive benchmark to systematically evaluate and study the reasoning abilities of SLMs trained from scratch or derived from LLMs through quantization, pruning, and distillation. We first establish a reliable evaluation criterion comparing available methods and LLM judges against our human evaluations. Then we present a study evaluating 72 diverse SLMs from six major model families across 17 reasoning benchmarks. We repeat all our experiments three times to ensure a robust assessment. Our findings show that: 1) reasoning ability in SLMs is strongly influenced by training methods and data quality rather than solely model scale; 2) quantization preserves reasoning capability, while pruning significantly disrupts it; 3) larger models consistently exhibit higher robustness against adversarial perturbations and intermediate reasoning, but certain smaller models closely match or exceed the larger models' performance. Our findings challenge the assumption that scaling is the only way to achieve strong reasoning. Instead, we foresee a future where SLMs with strong reasoning capabilities can be developed through structured training or post-training compression. Our THINKSLM Leaderboard is publicly available at: https://ctrlgaurav.github.io/thinkslm.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9283f6a7-8495-4bda-852d-fd5f0b256109Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMsXiangyu Wen, Junhua Huang, Zeju Li, Min Li et al.ICLR 2026 · 7 citations
- Do Large Language Models excel in Complex Logical Reasoning with Formal Language?Jin Jiang, Jianing Wang, Yuchen Yan, Yang Liu et al.EMNLP 2025 · 3 citations
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model ReasoningLi Wang, Changhao Zhang, Zengqi Xiu, Kai Lu et al.AAAI 2026 · 1 citation
- Enhancing Reasoning Abilities of Small LLMs with Cognitive AlignmentWenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang et al.EMNLP 2025
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question AnsweringFrancesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia et al.ACL 2026 · 1 citation
