Scaling Laws in Patchification: An Image Is Worth 50, 176 Tokens And More
Feng Wang, Yaodong Yu, Wei Shao, Yuyin Zhou, Alan L. Yuille, Cihang Xie
摘要
Since the introduction of Vision Transformer (ViT), patchification has long been regarded as a de facto image tokenization approach for plain visual architectures. By compressing the spatial size of images, this approach can effectively shorten the token sequence and reduce the computational cost of ViT-like plain architectures. In this work, we aim to thoroughly examine the information loss caused by this patchification-based compressive encoding paradigm and how it affects visual understanding. We conduct extensive patch size scaling experiments and excitedly observe an intriguing scaling law in patchification: the models can consistently benefit from decreased patch sizes and attain improved predictive performance, until it reaches the minimum patch size of 1×1, i.e., pixel tokenization. This conclusion is broadly applicable across different vision tasks, various input scales, and diverse architectures such as ViT and the recent Mamba models. Moreover, as a by-product, we discover that with smaller patches, task-specific decoder heads become less critical for dense prediction. In the experiments, we successfully scale up the visual sequence to an exceptional length of 50,176 tokens, achieving a competitive test accuracy of 84.6% with a base-sized model on the ImageNet-1k benchmark. We hope this study can provide insights and theoretical foundations for future works of building non-compressive vision models. Code is available at https://github.com/ wangf3014/Patch_Scaling .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a FewQishuai Wen, Zhiyuan Huang, Chun-Guang LiNeurIPS 2025 · 被引用 6 次
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language ModelsYangfu Li, Hongjian Zhan, Jiawei Chen, Yuning Gong 等CVPR 2026 · 被引用 6 次
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba ModelsTiming Yang, Feng Wang, Guoyizhe WeiCVPR 2026 · 被引用 2 次
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu 等ICML 2026 · 被引用 2 次
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal LearningXianhang Li, Yanqing Liu, Haoqin Tu, Cihang XieICCV 2025 · 被引用 1 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei 等CVPR 2025
- Faster Vision Transformers with Adaptive PatchesRohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang 等ICLR 2026 · 被引用 8 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou 等NeurIPS 2021 · 被引用 252 次
