Efficient Maximal Coding Rate Reduction by Variational Forms
Christina Baek, Ziyang Wu, Kwan Ho Ryan Chan, Tianjiao Ding, Yi Ma, Benjamin D. Haeffele
Abstract
The principle of Maximal Coding Rate Reduction (MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> ) has recently been proposed as a training objective for learning discriminative low-dimensional structures intrinsic to high-dimensional data to allow for more robust training than standard approaches, such as cross-entropy minimization. However, despite the advantages that have been shown for MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> training, MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> suffers from a significant computational cost due to the need to evaluate and differentiate a significant number of log-determinant terms that grows linearly with the number of classes. By taking advantage of variational forms of spectral functions of a matrix, we reformulate the MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> objective to a form that can scale significantly without compromising training accuracy. Experiments in image classification demonstrate that our proposed formulation results in a significant speed up over optimizing the original MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> objective directly and often results in higher quality learned representations. Further, our approach may be of independent interest in other models that require computation of log-determinant forms, such as in system identification or normalizing flow models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fafc66a4-daad-4050-8aa4-89acde04809eCited by top-tier papers2
- Image Clustering via the Principle of Rate Reduction in the Age of Pretrained ModelsTianzhe Chu, Shengbang Tong, Tianjiao Ding, Xili Dai et al.ICLR 2024 · 22 citations
- Unsupervised Manifold Linearizing and ClusteringTianjiao Ding, Shengbang Tong, Kwan Ho Ryan Chan, Xili Dai et al.ICCV 2023 · 19 citations
Builds on5
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song et al.NeurIPS 2020 · 265 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space PerspectiveLuca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli et al.ICLR 2022 · 67 citations
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
Related papers
- A Global Geometric Analysis of Maximal Coding Rate ReductionPeng Wang, Huikang Liu, Druv Pai, Yaodong Yu et al.ICML 2024 · 13 citations
- PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of Xiaojie Yu, Haibo Zhang, Jeremiah D. Deng, Lizhi PengICML 2026
- Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNetFengrong Li, Delin ChuICLR 2026
- Relative gradient optimization of the Jacobian term in unsupervised deep learningLuigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf et al.NeurIPS 2020 · 25 citations
- An In-depth Investigation of Sparse Rate Reduction in Transformer-like ModelsYunzhe Hu, Difan Zou, Dong XuNeurIPS 2024 · 4 citations
