Scaling White-Box Transformers for Vision
Jinrui Yang, Xianhang Li, Druv Pai, Yuyin Zhou, Yi Ma, Yaodong Yu, Cihang Xie
摘要
CRATE, a white-box transformer architecture designed to learn compressed and sparse representations, offers an intriguing alternative to standard vision transformers (ViTs) due to its inherent mathematical interpretability. Despite extensive investigations into the scaling behaviors of language and vision transformers, the scalability of CRATE remains an open question which this paper aims to address. Specifically, we propose CRATE-, featuring strategic yet minimal modifications to the sparse coding block in the CRATE architecture design, and a light training recipe designed to improve the scalability of CRATE. Through extensive experiments, we demonstrate that CRATE- can effectively scale with larger model sizes and datasets. For example, our CRATE--B substantially outperforms the prior best CRATE-B model accuracy on ImageNet classification by 3.7%, achieving an accuracy of 83.2%. Meanwhile, when scaling further, our CRATE--L obtains an ImageNet classification accuracy of 85.1%. More notably, these model performance improvements are achieved while preserving, and potentially even enhancing the interpretability of learned CRATE models, as we demonstrate through showing that the learned token representations of increasingly larger trained CRATE- models yield increasingly higher-quality unsupervised object segmentation of images. The project page is https://rayjryang.github.io/CRATE-alpha/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression PerspectiveQishuai Wen, Chun-Guang LiNeurIPS 2024 · 被引用 8 次
- Simplifying DINO via Coding Rate RegularizationZiyang Wu, Jingyuan Zhang, Druv Pai, Xudong Wang 等ICML 2025
- Towards White-Box Deep Wireless SensingXie Zhang, Yina Wang, Chenshu WuUbiComp 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu 等NeurIPS 2023 · 被引用 149 次
- Masked Completion via Structured Diffusion with White-Box TransformersDruv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu 等ICLR 2024 · 被引用 16 次
- Scaling Laws in Patchification: An Image Is Worth 50, 176 Tokens And MoreFeng Wang, Yaodong Yu, Wei Shao, Yuyin Zhou 等ICML 2025
- Bootstrapping SparseFormers from Vision Foundation ModelsZiteng Gao, Zhan Tong, Kevin Qinghong Lin, Joya Chen 等CVPR 2024
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song 等ICLR 2022 · 被引用 27 次
