Nuclear Norm Regularization for Deep Learning
Christopher Scarvelis, Justin M. Solomon
Abstract
Penalizing the nuclear norm of a function's Jacobian encourages it to locally behave like a low-rank linear map. Such functions vary locally along only a handful of directions, making the Jacobian nuclear norm a natural regularizer for machine learning problems. However, this regularizer is intractable for high-dimensional problems, as it requires computing a large Jacobian matrix and taking its singular value decomposition. We show how to efficiently penalize the Jacobian nuclear norm using techniques tailor-made for deep learning. We prove that for functions parametrized as compositions , one may equivalently penalize the average squared Frobenius norm of and . We then propose a denoising-style approximation that avoids the Jacobian computations altogether. Our method is simple, efficient, and accurate, enabling Jacobian nuclear norm regularization to scale to high-dimensional deep learning problems. We complement our theory with an empirical study of our regularizer's performance and investigate applications to denoising and representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector AlignmentNathanaël Haas, François Gatine, Augustin Cosse, Zied BouraouiICML 2026 · 1 citation
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
Builds on4
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Generalization in diffusion models arises from geometry-adaptive harmonic representationsZahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane MallatICLR 2024 · 168 citations
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 134 citations
Related papers
- Learning Representation from Neural Fisher Kernel with Low-rank ApproximationRuixiang Zhang, Shuangfei Zhai, Etai Littwin, Joshua M. SusskindICLR 2022 · 5 citations
- Operator SVD with Neural Networks via Nested Low-Rank ApproximationJongha Jon Ryu, Xiangxiang Xu, Hasan Sabri Melihcan Erol, Yuheng Bu et al.ICML 2024 · 10 citations
- Amortized Eigendecomposition for Neural NetworksTianbo Li, Zekun Shi, Jiaxi Zhao, Min LinNeurIPS 2024 · 4 citations
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- Fiedler Regularization: Learning Neural Networks with Graph SparsityEdric Tam, David B. DunsonICML 2020 · 15 citations
