On the Inductive Bias of Stacking Towards Improving Reasoning
Nikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi, Sashank Jakkam Reddi, Sanjiv Kumar
Abstract
Given the increasing scale of model sizes, novel training strategies like gradual stacking [Gong et al., 2019, Reddi et al., 2023] have garnered interest. Stacking enables efficient training by gradually growing the depth of a model in stages and using layers from a smaller model in an earlier stage to initialize the next stage. Although efficient for training, the model biases induced by such growing approaches are largely unexplored. In this work, we examine this fundamental aspect of gradual stacking, going beyond its efficiency benefits. We propose a variant of gradual stacking called MIDAS that can speed up language model training by up to 40%. Furthermore we discover an intriguing phenomenon: MIDAS is not only training-efficient but surprisingly also has an inductive bias towards improving downstream tasks, especially tasks that require reasoning abilities like reading comprehension and math problems, despite having similar or slightly worse perplexity compared to baseline training. To further analyze this inductive bias, we construct reasoning primitives -- simple synthetic tasks that are building blocks for reasoning -- and find that a model pretrained with stacking is significantly better than standard pretraining on these primitives, with and without fine-tuning. This provides stronger and more robust evidence for this inductive bias towards reasoning. These findings of training efficiency and inductive bias towards reasoning are verified at 1B, 2B and 8B parameter language models. Finally, we conjecture the underlying reason for this inductive bias by exploring the connection of stacking to looped models and provide strong supporting empirical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52f4691f-da00-48d4-9b91-58c790acf0e7Cited by top-tier papers11
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut ModulationAhmadreza Jeddi, Marco Ciccone, Babak TaatiICLR 2026 · 54 citations
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-ThoughtGuanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li et al.NeurIPS 2025 · 8 citations
- Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer ManifoldXinghan Li, Haodong Wen, Kaifeng LyuNeurIPS 2025 · 6 citations
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald et al.ICML 2026 · 2 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Looped Transformers as Programmable ComputersAngeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee et al.ICML 2023 · 175 citations
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- Understanding Contrastive Learning Requires Incorporating Inductive BiasesNikunj Saunshi, Jordan T. Ash, Surbhi Goel, Dipendra Misra et al.ICML 2022 · 130 citations
Related papers
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou et al.ICLR 2026 · 1 citation
- Curriculum-Guided Layer Scaling for Language Model PretrainingKaranpartap Singh, Neil Band, Ehsan AdeliICML 2026
- Efficient Training of Language Models using Few-Shot LearningSashank J. Reddi, Sobhan Miryoosefi, Stefani Karp, Shankar Krishnan et al.ICML 2023 · 22 citations
- Efficient stagewise pretraining via progressive subnetworksAbhishek Panigrahi, Nikunj Saunshi, Kaifeng Lyu, Sobhan Miryoosefi et al.ICLR 2025
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard et al.ICLR 2023 · 13 citations
