Transformer Layers as Painters
Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones
Abstract
Despite their nearly universal adoption for large language models, the internal workings of transformers are not well understood. We aim to better understand the impact of removing or reorganizing information throughout the layers of a pretrained transformer. Such an understanding could both yield better usage of existing models as well as to make architectural improvements to produce new variants. We present a series of empirical studies on frozen models that show that the lower and final layers of pretrained transformers differ from middle layers, but that middle layers have a surprising amount of uniformity. We further show that some classes of problems have robustness to skipping layers, running the layers in an order different from how they were trained, or running the layers in parallel. Our observations suggest that even frozen pretrained models may gracefully trade accuracy for latency by skipping layers or running layers in parallel.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd420cf-7d4c-4c4d-9a6b-a95752e99958Cited by top-tier papers22
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsDingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo et al.NeurIPS 2025 · 28 citations
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsJitai Hao, Hao Liu, Xinyan Xiao, Qiang Huang et al.ICLR 2026 · 18 citations
- Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization FailureBoshi Wang, Huan SunICLR 2026 · 16 citations
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen et al.NeurIPS 2025 · 16 citations
- The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language AbilitiesZixuan Qin, Qingchen Yu, Kunlin Lyu, Zhaoxin Fan et al.ICLR 2026 · 10 citations
Builds on4
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
- CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent LayersLongwei Zou, Qingyang Wang, Han Zhao, Jiangang Kong et al.ACL 2024
Related papers
- Remarkable Robustness of LLMs: Stages of Inference?Vedang Lad, Jin Hwa Lee, Wes Gurnee, Max TegmarkNeurIPS 2025 · 5 citations
- Frozen Pretrained Transformers as Universal Computation EnginesKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchAAAI 2022 · 133 citations
- The Hidden Space of Transformer Language AdaptersJesujoba Alabi, Marius Mosbach, Matan Eyal, Dietrich Klakow et al.ACL 2024
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMsZiyue Li, Yang Li, Tianyi ZhouICML 2026 · 4 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
