Lune

ICML2025Top-tier venue

LAuReL: Learned Augmented Residual Layer

Gaurav Menghani, Ravi Kumar, Sanjiv Kumar

2025Year
2Top-tier citations

Abstract

One of the core pillars of efficient deep learning methods are architectural improvements, such as residual/skip connections, which have led to significantly better model convergence and quality. Since their introduction, residual connections have become ubiquitous not only in convolutional neural networks but also in transformer-based architectures, the backbone of LLMs. In this paper, we introduce the Learned Augmented Residual Layer (LAUREL)-a novel generalization of the canonical residual connectiondesigned to serve as an in-situ replacement while outperforming it in both model quality and footprint metrics. Our experiments show that LAU-REL can enhance quality for both vision and language models while adding fewer parameters and incurring less latency and memory overhead than naively increasing parameter count. For example, on the ImageNet-1K task, LAU-REL achieves the same model quality improvements as naively adding an extra layer while using 2.6× fewer parameters. Similarly, when pretraining 1B and 4B parameter LLMs, LAUREL improves performance on a variety of challenging downstream evaluation tasks by 2.54% to 20.05%, while adding only 0.012% and 0.1% additional parameters, respectively.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7ee589de-719b-46e2-a641-275b3993e079

Cited by top-tier papers2

Ask how each one uses it

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines