White-Box Transformers via Sparse Rate Reduction
Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin D. Haeffele, Yi Ma
Abstract
In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces. The quality of the final representation can be measured by a unified objective function called sparse rate reduction. From this perspective, popular deep networks such as transformers can be naturally viewed as realizing iterative schemes to optimize this objective incrementally. Particularly, we show that the standard transformer block can be derived from alternating optimization on complementary parts of this objective: the multi-head self-attention operator can be viewed as a gradient descent step to compress the token sets by minimizing their lossy coding rate, and the subsequent multi-layer perceptron can be viewed as attempting to sparsify the representation of the tokens. This leads to a family of white-box transformer-like deep network architectures which are mathematically fully interpretable. Despite their simplicity, experiments show that these networks indeed learn to optimize the designed objective: they compress and sparsify representations of large-scale real-world vision datasets such as ImageNet, and achieve performance very close to thoroughly engineered transformers such as ViT. Code is at https://github.com/Ma-Lab-Berkeley/CRATE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa35fe8e-11d6-4bb2-a977-64e557159c8bCited by top-tier papers34
- U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationChenxin Li, Xinyu Liu, Wuyang Li, Cheng Wang et al.AAAI 2025 · 452 citations
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 25 citations
- Scaling White-Box Transformers for VisionJinrui Yang, Xianhang Li, Druv Pai, Yuyin Zhou et al.NeurIPS 2024 · 20 citations
- PaCE: Parsimonious Concept Engineering for Large Language ModelsJinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker et al.NeurIPS 2024 · 20 citations
- Arithmetic Feature Interaction Is Necessary for Deep Tabular LearningYi Cheng, Renjun Hu, Haochao Ying, Xing Shi et al.AAAI 2024 · 18 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Attention-Only Transformers via Unrolled Subspace DenoisingPeng Wang, Yifu Lu, Yaodong Yu, Druv Pai et al.ICML 2025
- Masked Completion via Structured Diffusion with White-Box TransformersDruv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu et al.ICLR 2024 · 16 citations
- An In-depth Investigation of Sparse Rate Reduction in Transformer-like ModelsYunzhe Hu, Difan Zou, Dong XuNeurIPS 2024 · 4 citations
- DiffRate : Differentiable Compression Rate for Efficient Vision TransformersMengzhao Chen, Wenqi Shao, Peng Xu, Mingbao Lin et al.ICCV 2023 · 87 citations
- A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample ComplexityHongkang Li, Meng Wang, Sijia Liu, Pin-Yu ChenICLR 2023 · 1 citation
