Lune

NeurIPS2025Top-tier venue

LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision

Anthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim, Daniel G. Kyrollos, James Green, Evan Shelhamer

2025Year
7Citations
1Top-tier citations

Abstract

Vision transformers are ever larger, more accurate, and more expensive to compute. The expense is even more extreme at high resolution as the number of tokens grows quadratically with the image size. We turn to adaptive computation to cope with this cost by learning to predict where to compute. Our LookWhere method divides the computation between a low-resolution selector and a high-resolution extractor without ever processing the full high-resolution input. We jointly pretrain the selector and extractor without task supervision by distillation from a selfsupervised teacher, in effect, learning where and what to compute simultaneously. Unlike prior token reduction methods, which pay to save by pruning alreadycomputed tokens, and prior token selection methods, which require complex and expensive per-task optimization, LookWhere economically and accurately selects and extracts transferrable representations of images. We show that LookWhere excels at sparse recognition on high-resolution inputs (Traffic Signs), maintaining accuracy while reducing FLOPs by up to 34× and time by 6×. It also excels at standard recognition tasks that are global (ImageNet classification) or local (ADE20K segmentation), improving accuracy while reducing time by 1.36×. See https://github.com/antofuller/lookwhere for the code and weights.

Teacher. We choose DINOv2 [1] as the teacher for its ViT architecture and internet-scale pretraining on diverse images without annotations. We use the variant with G=4 register tokens to reduce attention artifacts [35], as we use its attention to train the selector. The teacher processes inputs at resolution R high =518 with patch size P =14 for a grid of N high =37 patches. (This is unaltered from DINOv2.) To distill its attention, we extract the unnormalized attention among its patch tokens at the last layer and then average over queries and heads. This approximates where patch tokens contributed to the deepest teacher representation. To distill its representation, we extract the class and patch tokens at the last layer: z high ∈ R (1+N 2 high )×D . We never update the teacher: its parameters are fixed.

Losses and Updates. We train LookWhere jointly with three losses for what to compute, by distillation of the class token and patch tokens, and where to compute, by distillation of attention.

• Class Token Distillation: We distill the teacher's class token via mean-squared error (MSE): L cls = MSE(ẑ cls high , z cls high ). This trains the selector and extractor to learn global representations of the high-res input, given the low-res input and selected high-res patches.

• Patch Token Distillation: We distill the teacher's patch representations via the mean-squared error (MSE): L pat = MSE(ẑ pat high , z pat high ). This trains the selector and extractor to learn local representations of the high-res input, given the low-res input and selected high-res patches.

• Attention Distillation: We distill the teacher's attention map via the Kullback-Leibler (KL) divergence: L map = KL( Âhigh , A high ). This trains the selector to predict where the teacher computes.

Our pretraining loss is their sum L = λ cls L cls + λ pat L pat + λ map L map . We set λ cls , λ pat = 1, and λ map = 0.1. All losses are optimized end-to-end by the extractor and selector. We update the selector and extractor jointly and in parallel without custom batching, balancing, or tuning.

Initialization. We set the selector and extractor parameters to those of the teacher: DINOv2.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2d2e6ebb-d7fe-493b-9ec0-9acfe3abc99f

Cited by top-tier papers1

Ask how each one uses it

Builds on39

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines