LAuReL: Learned Augmented Residual Layer
Gaurav Menghani, Ravi Kumar, Sanjiv Kumar
摘要
One of the core pillars of efficient deep learning methods are architectural improvements, such as residual/skip connections, which have led to significantly better model convergence and quality. Since their introduction, residual connections have become ubiquitous not only in convolutional neural networks but also in transformer-based architectures, the backbone of LLMs. In this paper, we introduce the Learned Augmented Residual Layer (LAUREL)-a novel generalization of the canonical residual connectiondesigned to serve as an in-situ replacement while outperforming it in both model quality and footprint metrics. Our experiments show that LAU-REL can enhance quality for both vision and language models while adding fewer parameters and incurring less latency and memory overhead than naively increasing parameter count. For example, on the ImageNet-1K task, LAU-REL achieves the same model quality improvements as naively adding an extra layer while using 2.6× fewer parameters. Similarly, when pretraining 1B and 4B parameter LLMs, LAUREL improves performance on a variety of challenging downstream evaluation tasks by 2.54% to 20.05%, while adding only 0.012% and 0.1% additional parameters, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao 等ICML 2026 · 被引用 72 次
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng 等ICLR 2025
它引用的顶会 Paper5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Efficiency MisnomerMostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer 等ICLR 2022 · 被引用 116 次
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh 等NeurIPS 2023 · 被引用 13 次
相关 Paper
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang 等ICML 2026
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao 等ACL 2025 · 被引用 6 次
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou 等ICLR 2026 · 被引用 1 次
- LORS: Low-Rank Residual Structure for Parameter-Efficient Network StackingJialin Li, Qiang Nie, Weifu Fu, Yuhuan Lin 等CVPR 2024
