Vision Transformers Don't Need Trained Registers
Nick Jiang, Amil Dravid, Alexei A. Efros, Yossi Gandelsman
Abstract
We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers -the emergence of high-norm tokens that lead to noisy attention maps (Darcet et al., 2024). We observe that in multiple models (e.g., CLIP, DINOv2), a sparse set of neurons is responsible for concentrating high-norm activations on outlier tokens, leading to irregular attention patterns and degrading downstream visual processing. While the existing solution for removing these outliers involves retraining models from scratch with additional learned register tokens, we use our findings to create a training-free approach to mitigate these artifacts. By shifting the high-norm activations from our discovered register neurons into an additional untrained token, we can mimic the effect of register tokens on a model already trained without registers. We demonstrate that our method produces cleaner attention and feature maps, enhances performance over base models across multiple downstream visual tasks, and achieves results comparable to models explicitly trained with register tokens. We then extend test-time registers to off-the-shelf vision-language models, yielding cleaner attention-based, text-toimage attribution. Finally, we outline a simple mathematical model that reflects the observed behavior of register neurons and high norm tokens. Our results suggest that test-time registers effectively take on the role of register tokens at test-time, offering a training-free solution for any pre-trained model released without them. 1
• We demonstrate that activating the register neurons at test-time in other image locations shifts the high-norms to the corresponding tokens (Section 3.2).
• We present a training-free method for adding registers to models that were trained without them, by appending additional tokens and activating register neurons in their positions (Section 4).
• We evaluate the performance of models with test-time registers and show that it is comparable to models with trained registers, thus eliminating the need for retraining models with registers from scratch (Section 5).
Feature visualization in vision models. Visualizing features of computer vision models has been used for diagnostics long before the transition of the field to deep-learning (e.g., Vondrick et al. (2013)). Features in early CNN-based models were visualized to interpret their emergent computation (Zeiler
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6668298-ea23-4467-a3fe-01b977c0917dCited by top-tier papers10
- Vision Transformers with Self-Distilled RegistersZipeng Yan, Yinjie Chen, Chong Zhou, Bo Dai et al.NeurIPS 2025 · 17 citations
- The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language AbilitiesZixuan Qin, Qingchen Yu, Kunlin Lyu, Zhaoxin Fan et al.ICLR 2026 · 10 citations
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTsAlexander Lappe, Martin A. GieseNeurIPS 2025 · 9 citations
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2Joel Valdivia Ortega, Lorenz Lamm, Franziska Eckardt, Benedikt Schworm et al.NeurIPS 2025 · 1 citation
- Attention Sinks in Diffusion Transformers: A Causal AnalysisFANGZHENG WU, Brian SummaICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei et al.CVPR 2025
- DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object HallucinationXuan Gong, Tianshi Ming, Xinpeng Wang, Zhihua WeiEMNLP 2024 · 10 citations
- UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive RegisterCongpei Qiu, Zhaoyu Hu, Wei Ke, Zhuotao Tian et al.CVPR 2026
- Saliency-Driven Token Merging for Vision TransformersWeiying Xie, Xiaoyu Chen, Xin Zhang, Chenhe Hao et al.CVPR 2026
