An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
Yunzhe Hu, Difan Zou, Dong Xu
Abstract
Deep neural networks have long been criticized for being black-box. To unveil the inner workings of modern neural architectures, a recent work proposed an information-theoretic objective function called Sparse Rate Reduction (SRR) and interpreted its unrolled optimization as a Transformer-like model called Coding Rate Reduction Transformer (CRATE). However, the focus of the study was primarily on the basic implementation, and whether this objective is optimized in practice and its causal relationship to generalization remain elusive. Going beyond this study, we derive different implementations by analyzing layer-wise behaviors of CRATE, both theoretically and empirically. To reveal the predictive power of SRR on generalization, we collect a set of model variants induced by varied implementations and hyperparameters and evaluate SRR as a complexity measure based on its correlation with generalization. Surprisingly, we find out that SRR has a positive correlation coefficient and outperforms other baseline measures, such as path-norm and sharpness-based ones. Furthermore, we show that generalization can be improved using SRR as regularization on benchmark image classification datasets. We hope this paper can shed light on leveraging SRR to design principled models and study their generalization ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 9 citations
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a FewQishuai Wen, Zhiyuan Huang, Chun-Guang LiNeurIPS 2025 · 6 citations
- Hyper-SET: Designing Transformers via Hyperspherical Energy MinimizationYunzhe Hu, Difan Zou, Dong XuICLR 2026 · 2 citations
- PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of Xiaojie Yu, Haibo Zhang, Jeremiah D. Deng, Lizhi PengICML 2026
- Attention's forward pass and Frank-WolfeAlbert Alcalde, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026
Builds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
Related papers
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu et al.NeurIPS 2023 · 149 citations
- Towards Disentangling Information Paths with Coded ResNeXtApostolos Avranas, Marios KountourisNeurIPS 2022 · 1 citation
- Masked Completion via Structured Diffusion with White-Box TransformersDruv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu et al.ICLR 2024 · 16 citations
- Scaling White-Box Transformers for VisionJinrui Yang, Xianhang Li, Druv Pai, Yuyin Zhou et al.NeurIPS 2024 · 20 citations
- Towards White-Box Deep Wireless SensingXie Zhang, Yina Wang, Chenshu WuUbiComp 2026
