The Implicit Bias of Gradient Descent on Separable Multiclass Data
Hrithik Ravi, Clayton Scott, Daniel Soudry, Yutong Wang
摘要
Implicit bias describes the phenomenon where optimization-based training algorithms, without explicit regularization, show a preference for simple estimators even when more complex estimators have equal objective values. Multiple works have developed the theory of implicit bias for binary classification under the assumption that the loss satisfies an exponential tail property. However, there is a noticeable gap in analysis for multiclass classification, with only a handful of results which themselves are restricted to the cross-entropy loss. In this work, we employ the framework of Permutation Equivariant and Relative Margin-based (PERM) losses [Wang and Scott, 2024] to introduce a multiclass extension of the exponential tail property. This class of losses includes not only cross-entropy but also other losses. Using this framework, we extend the implicit bias result of Soudry et al. [2018] to multiclass classification. Furthermore, our proof techniques closely mirror those of the binary case, thus illustrating the power of the PERM framework for bridging the binary-multiclass gap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Do Neural Networks Need Gradient Descent to Generalize? A Theoretical StudyYotam Alexander, Yonatan Slutzky, Yuval Ran-Milo, Nadav CohenNeurIPS 2025 · 被引用 3 次
- Any-stepsize Gradient Descent for Separable Data under Fenchel-Young LossesHan Bao, Shinsaku Sakaue, Yuki TakezawaNeurIPS 2025 · 被引用 2 次
- The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean LabelsYonatan Slutzky, Yotam Alexander, Noam Razin, Nadav CohenNeurIPS 2025 · 被引用 2 次
- Variational Deep Learning via Implicit RegularizationJonathan Wenger, Beau Coker, Juraj Marusic, John Patrick CunninghamICLR 2026 · 被引用 1 次
- Post-Training with Policy Gradients: Optimality and the Base Model BarrierAlireza Mousavi-Hosseini, Murat ErdogduICML 2026 · 被引用 1 次
它引用的顶会 Paper3
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 被引用 94 次
- Fast margin maximization via dual accelerationZiwei Ji, Nathan Srebro, Matus TelgarskyICML 2021 · 被引用 42 次
相关 Paper
- Multiclass learning with margin: exponential rates with no bias-variance trade-offStefano Vigogna, Giacomo Meanti, Ernesto De Vito, Lorenzo RosascoICML 2022 · 被引用 3 次
- Temporal Imbalance of Positive and Negative Supervision in Class-Incremental LearningJinge Ma, Fengqing ZhuCVPR 2026 · 被引用 1 次
- Cut your Losses with SquentropyLike Hui, Mikhail Belkin, Stephen WrightICML 2023 · 被引用 9 次
- A Universal Growth Rate for Learning with Smooth Surrogate LossesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 被引用 27 次
- Implicit Bias of Spectal Descent and Muon on Multiclass Separable DataChen Fan, Mark Schmidt, Christos ThrampoulidisNeurIPS 2025
