Cut your Losses with Squentropy
Like Hui, Mikhail Belkin, Stephen Wright
Abstract
Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Recent work (Hui & Belkin, 2020) suggests that training using the (rescaled) square loss is often superior in terms of the classification accuracy. In this paper we propose the "squentropy" loss, which is the sum of two terms: the cross-entropy loss and the average square loss over the incorrect classes. We provide an extensive set of experiments on multi-class classification problems showing that the squentropy loss outperforms both the pure cross entropy and rescaled square losses in terms of the classification accuracy. We also demonstrate that it provides significantly better model calibration than either of these alternative losses and, furthermore, has less variance with respect to the random initialization. Additionally, in contrast to the square loss, squentropy loss can typically be trained using exactly the same optimization parameters, including the learning rate, as the standard crossentropy loss, making it a true "plug-and-play" replacement. Finally, unlike the rescaled square loss, multiclass squentropy contains no parameters that need to be adjusted.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi et al.NeurIPS 2024 · 7 citations
- Rethinking Confidence Scores and Thresholds in Pseudolabeling-based SSLHarit Vishwakarma, Yi Chen, Satya Sai Srinath Namburi GNVV, Sui Jiet Tay et al.ICML 2025
Builds on4
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 182 citations
- The Devil is in the Margin: Margin-based Label Smoothing for Network CalibrationBingyuan Liu, Ismail Ben Ayed, Adrian Galdran, Jose DolzCVPR 2022 · 61 citations
Related papers
- Understanding Square Loss in Training Overparametrized Neural Network ClassifiersTianyang Hu, Jun Wang, Wenjia Wang, Zhenguo LiNeurIPS 2022 · 23 citations
- PolyLoss: A Polynomial Expansion Perspective of Classification Loss FunctionsZhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk et al.ICLR 2022 · 189 citations
- Calibrating Deep Neural Networks by Pairwise ConstraintsJiacheng Cheng, Nuno VasconcelosCVPR 2022 · 21 citations
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu et al.NeurIPS 2022 · 93 citations
- Beyond calibration: estimating the grouping loss of modern neural networksAlexandre Perez-Lebel, Marine Le Morvan, Gaël VaroquauxICLR 2023 · 5 citations
