Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks
Like Hui, Mikhail Belkin
Abstract
Modern neural architectures for classification tasks are trained using the cross-entropy loss, which is widely believed to be empirically superior to the square loss. In this work we provide evidence indicating that this belief may not be well-founded. We explore several major neural architectures and a range of standard benchmark datasets for NLP, automatic speech recognition (ASR) and computer vision tasks to show that these architectures, with the same hyper-parameter settings as reported in the literature, perform comparably or better when trained with the square loss, even after equalizing computational resources. Indeed, we observe that the square loss produces better results in the dominant majority of NLP and ASR experiments. Cross-entropy appears to have a slight edge on computer vision tasks. We argue that there is little compelling empirical or theoretical evidence indicating a clear-cut advantage to the cross-entropy loss. Indeed, in our experiments, performance on nearly all non-vision tasks can be improved, sometimes significantly, by switching to the square loss. Furthermore, training with square loss appears to be less sensitive to the randomness in initialization. We posit that training using the square loss for classification needs to be a part of best practices of modern deep learning on equal footing with cross-entropy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers71
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville et al.NeurIPS 2021 · 378 citations
- Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding MethodsRandall Balestriero, Yann LeCunNeurIPS 2022 · 189 citations
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 182 citations
- Test Time Adaptation via Conjugate Pseudo-labelsSachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 152 citations
- Soft Calibration Objectives for Neural NetworksArchit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan et al.NeurIPS 2021 · 127 citations
Builds on1
Related papers
- Cut your Losses with SquentropyLike Hui, Mikhail Belkin, Stephen WrightICML 2023 · 9 citations
- Understanding Square Loss in Training Overparametrized Neural Network ClassifiersTianyang Hu, Jun Wang, Wenjia Wang, Zhenguo LiNeurIPS 2022 · 23 citations
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu et al.NeurIPS 2022 · 93 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You et al.ICML 2022 · 122 citations
