Lune

NeurIPS2025Top-tier venue

Vision Transformers Don't Need Trained Registers

Nick Jiang, Amil Dravid, Alexei A. Efros, Yossi Gandelsman

2025Year
50Citations
10Top-tier citations

Abstract

We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers -the emergence of high-norm tokens that lead to noisy attention maps (Darcet et al., 2024). We observe that in multiple models (e.g., CLIP, DINOv2), a sparse set of neurons is responsible for concentrating high-norm activations on outlier tokens, leading to irregular attention patterns and degrading downstream visual processing. While the existing solution for removing these outliers involves retraining models from scratch with additional learned register tokens, we use our findings to create a training-free approach to mitigate these artifacts. By shifting the high-norm activations from our discovered register neurons into an additional untrained token, we can mimic the effect of register tokens on a model already trained without registers. We demonstrate that our method produces cleaner attention and feature maps, enhances performance over base models across multiple downstream visual tasks, and achieves results comparable to models explicitly trained with register tokens. We then extend test-time registers to off-the-shelf vision-language models, yielding cleaner attention-based, text-toimage attribution. Finally, we outline a simple mathematical model that reflects the observed behavior of register neurons and high norm tokens. Our results suggest that test-time registers effectively take on the role of register tokens at test-time, offering a training-free solution for any pre-trained model released without them. 1

• We demonstrate that activating the register neurons at test-time in other image locations shifts the high-norms to the corresponding tokens (Section 3.2).

• We present a training-free method for adding registers to models that were trained without them, by appending additional tokens and activating register neurons in their positions (Section 4).

• We evaluate the performance of models with test-time registers and show that it is comparable to models with trained registers, thus eliminating the need for retraining models with registers from scratch (Section 5).

Feature visualization in vision models. Visualizing features of computer vision models has been used for diagnostics long before the transition of the field to deep-learning (e.g., Vondrick et al. (2013)). Features in early CNN-based models were visualized to interpret their emergent computation (Zeiler

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e6668298-ea23-4467-a3fe-01b977c0917d

Cited by top-tier papers10

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines