Lune

NeurIPS2025顶会

LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision

Anthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim, Daniel G. Kyrollos, James Green, Evan Shelhamer

2025年份
7被引次数
1顶会引用

摘要

Vision transformers are ever larger, more accurate, and more expensive to compute. The expense is even more extreme at high resolution as the number of tokens grows quadratically with the image size. We turn to adaptive computation to cope with this cost by learning to predict where to compute. Our LookWhere method divides the computation between a low-resolution selector and a high-resolution extractor without ever processing the full high-resolution input. We jointly pretrain the selector and extractor without task supervision by distillation from a selfsupervised teacher, in effect, learning where and what to compute simultaneously. Unlike prior token reduction methods, which pay to save by pruning alreadycomputed tokens, and prior token selection methods, which require complex and expensive per-task optimization, LookWhere economically and accurately selects and extracts transferrable representations of images. We show that LookWhere excels at sparse recognition on high-resolution inputs (Traffic Signs), maintaining accuracy while reducing FLOPs by up to 34× and time by 6×. It also excels at standard recognition tasks that are global (ImageNet classification) or local (ADE20K segmentation), improving accuracy while reducing time by 1.36×. See https://github.com/antofuller/lookwhere for the code and weights.

Teacher. We choose DINOv2 [1] as the teacher for its ViT architecture and internet-scale pretraining on diverse images without annotations. We use the variant with G=4 register tokens to reduce attention artifacts [35], as we use its attention to train the selector. The teacher processes inputs at resolution R high =518 with patch size P =14 for a grid of N high =37 patches. (This is unaltered from DINOv2.) To distill its attention, we extract the unnormalized attention among its patch tokens at the last layer and then average over queries and heads. This approximates where patch tokens contributed to the deepest teacher representation. To distill its representation, we extract the class and patch tokens at the last layer: z high ∈ R (1+N 2 high )×D . We never update the teacher: its parameters are fixed.

Losses and Updates. We train LookWhere jointly with three losses for what to compute, by distillation of the class token and patch tokens, and where to compute, by distillation of attention.

• Class Token Distillation: We distill the teacher's class token via mean-squared error (MSE): L cls = MSE(ẑ cls high , z cls high ). This trains the selector and extractor to learn global representations of the high-res input, given the low-res input and selected high-res patches.

• Patch Token Distillation: We distill the teacher's patch representations via the mean-squared error (MSE): L pat = MSE(ẑ pat high , z pat high ). This trains the selector and extractor to learn local representations of the high-res input, given the low-res input and selected high-res patches.

• Attention Distillation: We distill the teacher's attention map via the Kullback-Leibler (KL) divergence: L map = KL( Âhigh , A high ). This trains the selector to predict where the teacher computes.

Our pretraining loss is their sum L = λ cls L cls + λ pat L pat + λ map L map . We set λ cls , λ pat = 1, and λ map = 0.1. All losses are optimized end-to-end by the extractor and selector. We update the selector and extractor jointly and in parallel without custom batching, balancing, or tuning.

Initialization. We set the selector and extractor parameters to those of the teacher: DINOv2.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 2d2e6ebb-d7fe-493b-9ec0-9acfe3abc99f

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper39

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖