MLI Formula: A Nearly Scale-Invariant Solution with Noise Perturbation
Bowen Tao, Xin-Chun Li, De-Chuan Zhan
摘要
Monotonic Linear Interpolation (MLI) refers to the peculiar phenomenon that the error between the initial and converged model monotonically decreases along the linear interpolation, i.e., (1α)θ 0 + αθ F . Previous works focus on paired initial and converged points, relating MLI to the smoothness of the optimization trajectory. In this paper, we find a shocking fact that the error curves still exhibit a monotonic decrease when θ 0 is replaced with noise or even zero values, implying that the decreasing curve may be primarily related to the property of the converged model rather than the optimization trajectory. We further explore the relationship between αθ F and θ F and propose scale invariance properties in various cases, including Generalized Scale Invariance (GSI), Rectified Scale Invariance (RSI), and Normalized Scale Invariance (NSI). From an inverse perspective, the MLI formula is essentially an equation that adds varying levels of noise (i.e., (1 -α)ϵ) to a nearly scale-invariant network (i.e., αθ F ), resulting in a monotonically increasing error as the noise level rises. MLI is a special case where ϵ is equal to θ 0 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
相关 Paper
- On Monotonic Linear Interpolation of Neural Network ParametersJames Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort 等ICML 2021 · 被引用 11 次
- Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep NetworksXiang Wang, Annie N. Wang, Mo Zhou, Rong GeICLR 2023
- Stable Nonconvex-Nonconcave Training via Linear InterpolationThomas Pethick, Wanyun Xie, Volkan CevherNeurIPS 2023 · 被引用 8 次
- Divergence-Free Neural Networks with Application to Image DenoisingSébastien Herbreteau, Etienne MeunierICLR 2026
- Scalable Monotonic Neural NetworksHyunho Kim, Jong-Seok LeeICLR 2024 · 被引用 8 次
