Efficient Maximal Coding Rate Reduction by Variational Forms
Christina Baek, Ziyang Wu, Kwan Ho Ryan Chan, Tianjiao Ding, Yi Ma, Benjamin D. Haeffele
摘要
The principle of Maximal Coding Rate Reduction (MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> ) has recently been proposed as a training objective for learning discriminative low-dimensional structures intrinsic to high-dimensional data to allow for more robust training than standard approaches, such as cross-entropy minimization. However, despite the advantages that have been shown for MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> training, MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> suffers from a significant computational cost due to the need to evaluate and differentiate a significant number of log-determinant terms that grows linearly with the number of classes. By taking advantage of variational forms of spectral functions of a matrix, we reformulate the MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> objective to a form that can scale significantly without compromising training accuracy. Experiments in image classification demonstrate that our proposed formulation results in a significant speed up over optimizing the original MCR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> objective directly and often results in higher quality learned representations. Further, our approach may be of independent interest in other models that require computation of log-determinant forms, such as in system identification or normalizing flow models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Image Clustering via the Principle of Rate Reduction in the Age of Pretrained ModelsTianzhe Chu, Shengbang Tong, Tianjiao Ding, Xili Dai 等ICLR 2024 · 被引用 22 次
- Unsupervised Manifold Linearizing and ClusteringTianjiao Ding, Shengbang Tong, Kwan Ho Ryan Chan, Xili Dai 等ICCV 2023 · 被引用 19 次
它引用的顶会 Paper5
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song 等NeurIPS 2020 · 被引用 265 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space PerspectiveLuca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli 等ICLR 2022 · 被引用 67 次
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
相关 Paper
- A Global Geometric Analysis of Maximal Coding Rate ReductionPeng Wang, Huikang Liu, Druv Pai, Yaodong Yu 等ICML 2024 · 被引用 13 次
- PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of Xiaojie Yu, Haibo Zhang, Jeremiah D. Deng, Lizhi PengICML 2026
- Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNetFengrong Li, Delin ChuICLR 2026
- Relative gradient optimization of the Jacobian term in unsupervised deep learningLuigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf 等NeurIPS 2020 · 被引用 25 次
- An In-depth Investigation of Sparse Rate Reduction in Transformer-like ModelsYunzhe Hu, Difan Zou, Dong XuNeurIPS 2024 · 被引用 4 次
