An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
Duy-Kien Nguyen, Mido Assran, Unnat Jain, Martin R. Oswald, Cees G. M. Snoek, Xinlei Chen
摘要
This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a token and achieve highly performant results. This is substantially different from the popular design in Vision Transformer, which maintains the inductive bias from ConvNets towards local neighborhoods (e.g., by treating each 16x16 patch as a token). We showcase the effectiveness of pixels-as-tokens across three well-studied computer vision tasks: supervised learning for classification and regression, self-supervised learning via masked autoencoding, and image generation with diffusion models. Although it's computationally less practical to directly operate on individual pixels, we believe the community must be made aware of this surprising piece of knowledge when devising the next generation of neural network architectures for computer vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- NeurIPT: Foundation Model for Neural InterfacesZitao Fang, Chenxuan Li, Hongting Zhou, Shuyang Yu 等NeurIPS 2025 · 被引用 16 次
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a FewQishuai Wen, Zhiyuan Huang, Chun-Guang LiNeurIPS 2025 · 被引用 6 次
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image RetrievalSiting Li, Xiang Gao, Simon S. DuNeurIPS 2025 · 被引用 5 次
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu 等ICML 2026 · 被引用 2 次
- DeltaDorsal: Enhancing Hand Pose Estimation with Dorsal Features in Egocentric ViewsWilliam Huang, Siyou Pei, Leyi Zou, Eric J. Gonzalez 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- Differentiable Hierarchical Visual TokenizationMarius Aasan, Martine Hjelkrem-Tan, Nico Catalano, Changkyu Choi 等NeurIPS 2025 · 被引用 4 次
- UniNeXt: Exploring A Unified Architecture for Vision RecognitionFangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang 等ACM MM 2023 · 被引用 15 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
