Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Ziyue Li, Yang Li, Tianyi Zhou
Abstract
Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic "program-of-layers (PoLar)", where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM’s latent reasoning capacity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6d9122f-f09a-40c3-a6ca-e88a04490ce3Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim et al.NeurIPS 2025 · 143 citations
- Faster Depth-Adaptive TransformersYijin Liu, Fandong Meng, Jie Zhou, Yufeng Chen et al.AAAI 2021 · 44 citations
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani et al.NeurIPS 2024 · 29 citations
Related papers
- Dr.LLM: Dynamic Layer Routing in LLMsAhmed Heakl, Martin Gubri, Salman Khan, Sangdoo Yun et al.ICLR 2026 · 11 citations
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao et al.NeurIPS 2024 · 39 citations
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
- Rethinking LLM Reasoning: From Explicit Trajectories to Latent RepresentationsCong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen et al.ICLR 2026
- Latent Reasoning in TRMs is Secretly a Policy Improvement OperatorArip Asadulaev, Rayan Banerjee, Fakhri Karray, Martin TakacICML 2026 · 4 citations
