Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Ziyue Li, Yang Li, Tianyi Zhou
摘要
Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic "program-of-layers (PoLar)", where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM’s latent reasoning capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim 等NeurIPS 2025 · 被引用 143 次
- Faster Depth-Adaptive TransformersYijin Liu, Fandong Meng, Jie Zhou, Yufeng Chen 等AAAI 2021 · 被引用 44 次
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani 等NeurIPS 2024 · 被引用 29 次
相关 Paper
- Dr.LLM: Dynamic Layer Routing in LLMsAhmed Heakl, Martin Gubri, Salman Khan, Sangdoo Yun 等ICLR 2026 · 被引用 11 次
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao 等NeurIPS 2024 · 被引用 39 次
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 被引用 110 次
- Rethinking LLM Reasoning: From Explicit Trajectories to Latent RepresentationsCong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen 等ICLR 2026
- Latent Reasoning in TRMs is Secretly a Policy Improvement OperatorArip Asadulaev, Rayan Banerjee, Fakhri Karray, Martin TakacICML 2026 · 被引用 4 次
