LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
Anthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim, Daniel G. Kyrollos, James Green, Evan Shelhamer
Abstract
Vision transformers are ever larger, more accurate, and more expensive to compute. The expense is even more extreme at high resolution as the number of tokens grows quadratically with the image size. We turn to adaptive computation to cope with this cost by learning to predict where to compute. Our LookWhere method divides the computation between a low-resolution selector and a high-resolution extractor without ever processing the full high-resolution input. We jointly pretrain the selector and extractor without task supervision by distillation from a selfsupervised teacher, in effect, learning where and what to compute simultaneously. Unlike prior token reduction methods, which pay to save by pruning alreadycomputed tokens, and prior token selection methods, which require complex and expensive per-task optimization, LookWhere economically and accurately selects and extracts transferrable representations of images. We show that LookWhere excels at sparse recognition on high-resolution inputs (Traffic Signs), maintaining accuracy while reducing FLOPs by up to 34× and time by 6×. It also excels at standard recognition tasks that are global (ImageNet classification) or local (ADE20K segmentation), improving accuracy while reducing time by 1.36×. See https://github.com/antofuller/lookwhere for the code and weights.
Teacher. We choose DINOv2 [1] as the teacher for its ViT architecture and internet-scale pretraining on diverse images without annotations. We use the variant with G=4 register tokens to reduce attention artifacts [35], as we use its attention to train the selector. The teacher processes inputs at resolution R high =518 with patch size P =14 for a grid of N high =37 patches. (This is unaltered from DINOv2.) To distill its attention, we extract the unnormalized attention among its patch tokens at the last layer and then average over queries and heads. This approximates where patch tokens contributed to the deepest teacher representation. To distill its representation, we extract the class and patch tokens at the last layer: z high ∈ R (1+N 2 high )×D . We never update the teacher: its parameters are fixed.
Losses and Updates. We train LookWhere jointly with three losses for what to compute, by distillation of the class token and patch tokens, and where to compute, by distillation of attention.
• Class Token Distillation: We distill the teacher's class token via mean-squared error (MSE): L cls = MSE(ẑ cls high , z cls high ). This trains the selector and extractor to learn global representations of the high-res input, given the low-res input and selected high-res patches.
• Patch Token Distillation: We distill the teacher's patch representations via the mean-squared error (MSE): L pat = MSE(ẑ pat high , z pat high ). This trains the selector and extractor to learn local representations of the high-res input, given the low-res input and selected high-res patches.
• Attention Distillation: We distill the teacher's attention map via the Kullback-Leibler (KL) divergence: L map = KL( Âhigh , A high ). This trains the selector to predict where the teacher computes.
Our pretraining loss is their sum L = λ cls L cls + λ pat L pat + λ map L map . We set λ cls , λ pat = 1, and λ map = 0.1. All losses are optimized end-to-end by the extractor and selector. We update the selector and extractor jointly and in parallel without custom batching, balancing, or tuning.
Initialization. We set the selector and extractor parameters to those of the teacher: DINOv2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d2e6ebb-d7fe-493b-9ec0-9acfe3abc99fCited by top-tier papers1
Ask how each one uses itBuilds on39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuningWeicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo et al.NeurIPS 2022 · 52 citations
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya et al.CVPR 2022 · 288 citations
- LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image RecognitionYoubing Hu, Yun Cheng, Anqi Lu, Zhiqiang Cao et al.AAAI 2024 · 25 citations
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision TransformersCong Wei, Brendan Duke, Ruowei Jiang, Parham Aarabi et al.CVPR 2023
