Vision Transformers Need More Than Registers
Cheng Shi, Yizhou Yu, Sibei Yang
摘要
Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, artifacts in ViTs are widely observed across different supervision paradigms and downstream tasks. Through systematic analysis of artifacts in ViTs, we find that their fundamental mechanisms have yet to be sufficiently elucidated. In this paper, through systematic analysis, we conclude that these artifacts originate from a lazy aggregation behavior: ViT uses semantically irrelevant background patches as shortcuts to represent global semantics, driven by global attention and Coarse-grained semantic supervision. Our solution selectively integrates patch features into the CLS token, reducing the influence of background-dominated shortcuts and consistently improving performance across 12 benchmarks under label-, text-, and self-supervision. We hope this work offers a new perspective on ViT behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Parameters as Experts: Adapting Vision Models with Dynamic Parameter RoutingMeng Lou, Stanley Yu, Yizhou YuICML 2026 · 被引用 1 次
- VGGT-ΩJianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev 等CVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 被引用 52 次
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTsAlexander Lappe, Martin A. GieseNeurIPS 2025 · 被引用 9 次
- BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye 等ICCV 2023 · 被引用 9 次
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang 等AAAI 2025 · 被引用 14 次
