LAuReL: Learned Augmented Residual Layer
Gaurav Menghani, Ravi Kumar, Sanjiv Kumar
Abstract
One of the core pillars of efficient deep learning methods are architectural improvements, such as residual/skip connections, which have led to significantly better model convergence and quality. Since their introduction, residual connections have become ubiquitous not only in convolutional neural networks but also in transformer-based architectures, the backbone of LLMs. In this paper, we introduce the Learned Augmented Residual Layer (LAUREL)-a novel generalization of the canonical residual connectiondesigned to serve as an in-situ replacement while outperforming it in both model quality and footprint metrics. Our experiments show that LAU-REL can enhance quality for both vision and language models while adding fewer parameters and incurring less latency and memory overhead than naively increasing parameter count. For example, on the ImageNet-1K task, LAU-REL achieves the same model quality improvements as naively adding an extra layer while using 2.6× fewer parameters. Similarly, when pretraining 1B and 4B parameter LLMs, LAUREL improves performance on a variety of challenging downstream evaluation tasks by 2.54% to 20.05%, while adding only 0.012% and 0.1% additional parameters, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ee589de-719b-46e2-a641-275b3993e079Cited by top-tier papers2
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao et al.ICML 2026 · 72 citations
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng et al.ICLR 2025
Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Efficiency MisnomerMostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer et al.ICLR 2022 · 116 citations
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang et al.ICLR 2023 · 52 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh et al.NeurIPS 2023 · 13 citations
Related papers
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang et al.ICML 2026
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao et al.ACL 2025 · 6 citations
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou et al.ICLR 2026 · 1 citation
- LORS: Low-Rank Residual Structure for Parameter-Efficient Network StackingJialin Li, Qiang Nie, Weifu Fu, Yuhuan Lin et al.CVPR 2024
