Mimetic Initialization of Self-Attention Layers
Asher Trockman, J. Zico Kolter
摘要
It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they "look" more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique mimetic initialization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionRomain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux 等ICML 2024 · 被引用 62 次
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 被引用 52 次
- What can a Single Attention Layer Learn? A Study Through the Random Features LensHengyu Fu, Tianyu Guo, Yu Bai, Song MeiNeurIPS 2023 · 被引用 47 次
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin 等ICLR 2024 · 被引用 40 次
- In-Context Learning with Transformers: Softmax Attention Adapts to Function LipschitznessLiam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi 等NeurIPS 2024 · 被引用 33 次
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- Conditioned Initialization for AttentionHemanth Saratchandran, Simon LuceyICLR 2026
- On the Surprising Effectiveness of Attention Transfer for Vision TransformersAlexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak 等NeurIPS 2024 · 被引用 21 次
- Structured Initialization for Vision TransformersJianqiao Zheng, Xueqian Li, Hemanth Saratchandran, Simon LuceyNeurIPS 2025 · 被引用 6 次
- Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?Boris Knyazev, Doha Hwang, Simon Lacoste-JulienICML 2023 · 被引用 31 次
- Can You Learn to See Without Images? Procedural Warm-Up for Vision TransformersZachary Shinnick, Liangze Jiang, Hemanth Saratchandran, Damien Teney 等CVPR 2026 · 被引用 9 次
