Value Residual Learning
Zhanchao Zhou, Tianyi Wu, Zhiyun Jiang, Fares Obeid, Zhenzhong Lan
Abstract
While Transformer models have achieved remarkable success in various domains, the effectiveness of information propagation through deep networks remains a critical challenge. Standard hidden state residuals often fail to adequately preserve initial token-level information in deeper layers. This paper introduces ResFormer, a novel architecture that enhances information flow by incorporating value residual connections in addition to hidden state residuals. And a variant is SVFormer, where all layers share the first layer's value embedding. Comprehensive empirical evidence demonstrates ResFormer achieves equivalent validation loss with 16.11% fewer model parameters and 20.3% less training data compared to Transformer, while maintaining similar memory usage and computational cost. Besides, SVFormer reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KV-efficient methods, yielding further reductions in KV cache, with performance influenced by sequence length and cumulative learning rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aba47c60-1a44-4785-bfbd-3223a6423c48Cited by top-tier papers4
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationFrançois Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe et al.NeurIPS 2025 · 23 citations
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value HeadsZhoutong Wu, Yuan Zhang, Yiming Dong, Chenheng Zhang et al.NeurIPS 2025 · 4 citations
- MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language GuidanceZihan Cao, Yu Zhong, Ziqi Wang, Liang-Jian DengICCV 2025 · 1 citation
- Attention Projection Mixing with Exogenous AnchorsJonathan SuICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
- ResFormer: All-Time Reservoir Memory for Long Sequence ClassificationHongbo Liu, Jia XuEMNLP 2025
- Residual Matrix Transformers: Scaling the Size of the Residual StreamBrian Mak, Jeffrey FlaniganICML 2025
- SRFormer: Permuted Self-Attention for Single Image Super-ResolutionYupeng Zhou, Zhen Li, Chun-Le Guo, Song Bai et al.ICCV 2023
- GRFormer: Grouped Residual Self-Attention for Lightweight Single Image Super-ResolutionYuzhen Li, Zehang Deng, Yuxin Cao, Lihua LiuACM MM 2024 · 9 citations
