Understanding Square Loss in Training Overparametrized Neural Network Classifiers
Tianyang Hu, Jun Wang, Wenjia Wang, Zhenguo Li
Abstract
Deep learning has achieved many breakthroughs in modern classification tasks. Numerous architectures have been proposed for different data structures but when it comes to the loss function, the cross-entropy loss is the predominant choice. Recently, several alternative losses have seen revived interests for deep classifiers. In particular, empirical evidence seems to promote square loss but a theoretical justification is still lacking. In this work, we contribute to the theoretical understanding of square loss in classification by systematically investigating how it performs for overparametrized neural networks in the neural tangent kernel (NTK) regime. Interesting properties regarding the generalization error, robustness, and calibration error are revealed. We consider two cases, according to whether classes are separable or not. In the general non-separable case, fast convergence rate is established for both misclassification rate and calibration error. When classes are separable, the misclassification rate improves to be exponentially fast. Further, the resulting margin is proven to be lower bounded away from zero, providing theoretical guarantees for robustness. We expect our findings to hold beyond the NTK regime and translate to practical settings. To this end, we conduct extensive empirical studies on practical neural networks, demonstrating the effectiveness of square loss in both synthetic low-dimensional data and real image data. Comparing to cross-entropy, square loss has comparable generalization error but noticeable advantages in robustness and model calibration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fe254e1-95e0-4f2b-9c66-343cc22d4b1cCited by top-tier papers7
- From Tempered to Benign Overfitting in ReLU Neural NetworksGuy Kornowski, Gilad Yehudai, Ohad ShamirNeurIPS 2023 · 18 citations
- Explore and Exploit the Diverse Knowledge in Model Zoo for Domain GeneralizationYimeng Chen, Tianyang Hu, Fengwei Zhou, Zhenguo Li et al.ICML 2023 · 14 citations
- Robust Classification via Regression for Learning with Noisy LabelsErik Englesson, Hossein AzizpourICLR 2024 · 12 citations
- GraphCleaner: Detecting Mislabelled Samples in Popular Graph Learning BenchmarksYuwen Li, Miao Xiong, Bryan HooiICML 2023 · 10 citations
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos et al.ICLR 2024 · 2 citations
Builds on15
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- MMA Training: Direct Input Space Margin Maximization through Adversarial TrainingGavin Weiguang Ding, Yash Sharma, Kry Yik Chau Lui, Ruitong HuangICLR 2020 · 308 citations
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song et al.NeurIPS 2020 · 265 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
Related papers
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
- Cut your Losses with SquentropyLike Hui, Mikhail Belkin, Stephen WrightICML 2023 · 9 citations
- Divergence of Neural Tangent Kernel in Classification ProblemsZixiong Yu, Songtao Tian, Guhan ChenICLR 2025
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning AnalysisBaihe Huang, Xiaoxiao Li, Zhao Song, Xin YangICML 2021 · 66 citations
