TrAct: Making First-layer Pre-Activations Trainable
Felix Petersen, Christian Borgelt, Stefano Ermon
Abstract
We consider the training of the first layer of vision models and notice the clear relationship between pixel values and gradient update magnitudes: the gradients arriving at the weights of a first layer are by definition directly proportional to (normalized) input pixel values. Thus, an image with low contrast has a smaller impact on learning than an image with higher contrast, and a very bright or very dark image has a stronger impact on the weights than an image with moderate brightness. In this work, we propose performing gradient descent on the embeddings produced by the first layer of the model. However, switching to discrete inputs with an embedding layer is not a reasonable option for vision models. Thus, we propose the conceptual procedure of (i) a gradient descent step on first layer activations to construct an activation proposal, and (ii) finding the optimal weights of the first layer, i.e., those weights which minimize the squared distance to the activation proposal. We provide a closed form solution of the procedure and adjust it for robust stochastic training while computing everything efficiently. Empirically, we find that TrAct (Training Activations) speeds up training by factors between 1.25x and 4x while requiring only a small computational overhead. We demonstrate the utility of TrAct with different optimizers for a range of different vision models including convolutional and transformer architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23aeae69-7cad-4da5-866c-e5b9e883e9cfBuilds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 114 citations
- Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent GeneralizationNeha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer et al.ICML 2021 · 18 citations
- Newton Losses: Using Curvature Information for Learning with Differentiable AlgorithmsFelix Petersen, Christian Borgelt, Tobias Sutter, Hilde Kuehne et al.NeurIPS 2024 · 3 citations
Related papers
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 14 citations
- Budgeted Training for Vision TransformerZhuofan Xia, Xuran Pan, Xuan Jin, Yuan He et al.ICLR 2023
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 10 citations
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick et al.ICLR 2022 · 47 citations
