An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
Yunzhe Hu, Difan Zou, Dong Xu
摘要
Deep neural networks have long been criticized for being black-box. To unveil the inner workings of modern neural architectures, a recent work proposed an information-theoretic objective function called Sparse Rate Reduction (SRR) and interpreted its unrolled optimization as a Transformer-like model called Coding Rate Reduction Transformer (CRATE). However, the focus of the study was primarily on the basic implementation, and whether this objective is optimized in practice and its causal relationship to generalization remain elusive. Going beyond this study, we derive different implementations by analyzing layer-wise behaviors of CRATE, both theoretically and empirically. To reveal the predictive power of SRR on generalization, we collect a set of model variants induced by varied implementations and hyperparameters and evaluate SRR as a complexity measure based on its correlation with generalization. Surprisingly, we find out that SRR has a positive correlation coefficient and outperforms other baseline measures, such as path-norm and sharpness-based ones. Furthermore, we show that generalization can be improved using SRR as regularization on benchmark image classification datasets. We hope this paper can shed light on leveraging SRR to design principled models and study their generalization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 被引用 9 次
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a FewQishuai Wen, Zhiyuan Huang, Chun-Guang LiNeurIPS 2025 · 被引用 6 次
- Hyper-SET: Designing Transformers via Hyperspherical Energy MinimizationYunzhe Hu, Difan Zou, Dong XuICLR 2026 · 被引用 2 次
- PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of Xiaojie Yu, Haibo Zhang, Jeremiah D. Deng, Lizhi PengICML 2026
- Attention's forward pass and Frank-WolfeAlbert Alcalde, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026
它引用的顶会 Paper24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl 等ICLR 2021 · 被引用 620 次
相关 Paper
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu 等NeurIPS 2023 · 被引用 149 次
- Towards Disentangling Information Paths with Coded ResNeXtApostolos Avranas, Marios KountourisNeurIPS 2022 · 被引用 1 次
- Masked Completion via Structured Diffusion with White-Box TransformersDruv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu 等ICLR 2024 · 被引用 16 次
- Scaling White-Box Transformers for VisionJinrui Yang, Xianhang Li, Druv Pai, Yuyin Zhou 等NeurIPS 2024 · 被引用 20 次
- Towards White-Box Deep Wireless SensingXie Zhang, Yina Wang, Chenshu WuUbiComp 2026
