Information-Theoretic Local Minima Characterization and Regularization
Zhiwei Jia, Hao Su
摘要
Recent advances in deep learning theory have evoked the study of generalizability across different local minima of deep neural networks (DNNs). While current work focused on either discovering properties of good local minima or developing regularization techniques to induce good local minima, no approach exists that can tackle both problems. We achieve these two goals successfully in a unified manner. Specifically, based on the observed Fisher information we propose a metric both strongly indicative of generalizability of local minima and effectively applied as a practical regularizer. We provide theoretical analysis including a generalization bound and empirically demonstrate the success of our approach in both capturing and improving the generalizability of DNNs. Experiments are performed on CIFAR-10, CIFAR-100 and ImageNet for various network architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Flatness-Aware Minimization for Domain GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Yancheng Dong 等ICCV 2023 · 被引用 37 次
- Semantically Robust Unpaired Image Translation for Data with Unmatched Semantics StatisticsZhiwei Jia, Bodi Yuan, Kangkang Wang, Hong Wu 等ICCV 2021 · 被引用 26 次
- Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit BiasRyo Karakida, Tomoumi Takase, Tomohiro Hayase, Kazuki OsawaICML 2023 · 被引用 23 次
- Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and FlatnessMingyuan Fan, Xiaodan Li, Cen Chen, Wenmeng Zhou 等NeurIPS 2024 · 被引用 13 次
- The Hessian perspective into the Nature of Convolutional Neural NetworksSidak Pal Singh, Thomas Hofmann, Bernhard SchölkopfICML 2023 · 被引用 12 次
相关 Paper
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg 等ICML 2021 · 被引用 78 次
- Inconsistency-Aware Minimization: Improving Generalization with Unlabeled DataHee-Sung Kim, Hyeonseong Kim, Sungyoon LeeICML 2026
- Inverse-Reference Priors for Fisher Regularization of Bayesian Neural NetworksKeunseo Kim, Eun-Yeol Ma, Jeongman Choi, Heeyoung KimAAAI 2023 · 被引用 2 次
- Embedding Principle of Loss Landscape of Deep Neural NetworksYaoyu Zhang, Zhongwang Zhang, Tao Luo, Zhi-Qin John XuNeurIPS 2021 · 被引用 48 次
- Regularizing Neural Networks via Adversarial Model PerturbationYaowei Zheng, Richong Zhang, Yongyi MaoCVPR 2021
