Hard labels sampled from sparse targets mislead rotation invariant algorithms
Avrajit Ghosh, Bin Yu, Manfred Warmuth, Peter Bartlett
摘要
One of the most common machine learning setups is logistic regression. In many classification models, including neural networks, the final prediction is obtained by applying a logistic link function to a linear score. In binary logistic regression, the feedback can be either soft labels, corresponding to the true conditional probability of the data (as in distillation), or sampled hard labels (taking values ). We point out a fundamental problem that arises even in a particularly favorable setting, where the goal is to learn a noise-free soft target of the form . In the over-constrained case (i.e. the number of samples exceeds the input dimension ) with examples , it is sufficient to recover and hence achieve the Bayes risk. However, we prove that when the examples are labeled by hard labels sampled from the same conditional distribution and is -sparse, then rotation-invariant algorithms are provably suboptimal: they incur an excess risk , while there are simple non-rotation invariant algorithms with excess risk . The simplest rotation invariant algorithm is gradient descent on the logistic loss (with early stopping). A simple non-rotation-invariant algorithm for sparse targets that achieves the above upper bounds uses gradient descent on the weights , where now the linear weight is reparameterized as .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- The Implicit Bias of Depth: How Incremental Learning Drives GeneralizationDaniel Gissin, Shai Shalev-Shwartz, Amit DanielyICLR 2020 · 被引用 90 次
- Saddle-to-Saddle Dynamics in Diagonal Linear NetworksScott Pesme, Nicolas FlammarionNeurIPS 2023 · 被引用 68 次
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 被引用 57 次
相关 Paper
- Early-stopped neural networks are consistentZiwei Ji, Justin D. Li, Matus TelgarskyNeurIPS 2021 · 被引用 58 次
- Sparse Linear Regression Is Easy on Random SupportsGautam Chandrasekaran, Raghu Meka, Konstantinos StavropoulosSTOC 2026
- A Precise Performance Analysis of Support Vector RegressionHoussem Sifaou, Abla Kammoun, Mohamed-Slim AlouiniICML 2021 · 被引用 8 次
- Agnostic Learning of a Single Neuron with Gradient DescentSpencer Frei, Yuan Cao, Quanquan GuNeurIPS 2020 · 被引用 68 次
- Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic RegressionJingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin YuICML 2025
