Value Residual Learning
Zhanchao Zhou, Tianyi Wu, Zhiyun Jiang, Fares Obeid, Zhenzhong Lan
摘要
While Transformer models have achieved remarkable success in various domains, the effectiveness of information propagation through deep networks remains a critical challenge. Standard hidden state residuals often fail to adequately preserve initial token-level information in deeper layers. This paper introduces ResFormer, a novel architecture that enhances information flow by incorporating value residual connections in addition to hidden state residuals. And a variant is SVFormer, where all layers share the first layer's value embedding. Comprehensive empirical evidence demonstrates ResFormer achieves equivalent validation loss with 16.11% fewer model parameters and 20.3% less training data compared to Transformer, while maintaining similar memory usage and computational cost. Besides, SVFormer reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KV-efficient methods, yielding further reductions in KV cache, with performance influenced by sequence length and cumulative learning rate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationFrançois Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe 等NeurIPS 2025 · 被引用 23 次
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value HeadsZhoutong Wu, Yuan Zhang, Yiming Dong, Chenheng Zhang 等NeurIPS 2025 · 被引用 4 次
- MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language GuidanceZihan Cao, Yu Zhong, Ziqi Wang, Liang-Jian DengICCV 2025 · 被引用 1 次
- Attention Projection Mixing with Exogenous AnchorsJonathan SuICML 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- ResFormer: All-Time Reservoir Memory for Long Sequence ClassificationHongbo Liu, Jia XuEMNLP 2025
- Residual Matrix Transformers: Scaling the Size of the Residual StreamBrian Mak, Jeffrey FlaniganICML 2025
- SRFormer: Permuted Self-Attention for Single Image Super-ResolutionYupeng Zhou, Zhen Li, Chun-Le Guo, Song Bai 等ICCV 2023
- GRFormer: Grouped Residual Self-Attention for Lightweight Single Image Super-ResolutionYuzhen Li, Zehang Deng, Yuxin Cao, Lihua LiuACM MM 2024 · 被引用 9 次
