Supervised learning: no loss no cry
Richard Nock, Aditya Krishna Menon
Abstract
Supervised learning requires the specification of a loss function to minimise. While the theory of admissible losses from both a computational and statistical perspective is well-developed, these offer a panoply of different choices. In practice, this choice is typically made in an ad hoc manner. In hopes of making this procedure more principled, the problem of learning the loss function for a downstream task (e.g., classification) has garnered recent interest. However, works in this area have been generally empirical in nature. In this paper, we revisit the SLIsotron algorithm of Kakade et al. (2011) through a novel lens, derive a generalisation based on Bregman divergences, and show how it provides a principled procedure for learning the loss. In detail, we cast SLIsotron as learning a loss from a family of composite square losses. By interpreting this through the lens of proper losses, we derive a generalisation of SLIsotron based on Bregman divergences. The resulting BregmanTron algorithm jointly learns the loss along with the classifier. It comes equipped with a simple guarantee of convergence for the loss it learns, and its set of possible outputs comes with a guarantee of agnostic approximability of Bayes rule. Experiments indicate that the BregmanTron substantially outperforms the SLIsotron, and that the loss it learns can be minimized by other algorithms for different tasks, thereby opening the interesting problem of loss transfer between domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- All your loss are belong to BayesChristian J. Walder, Richard NockNeurIPS 2020 · 6 citations
- Random Classification Noise does not defeat All Convex Potential Boosters Irrespective of Model ChoiceYishay Mansour, Richard Nock, Robert C. WilliamsonICML 2023 · 4 citations
- Hyperbolic Embeddings of Supervised ModelsRichard Nock, Ehsan Amid, Frank Nielsen, Alexander Soen et al.NeurIPS 2024 · 2 citations
- LegendreTron: Uprising Proper Multiclass Loss LearningKevin H. Lam, Christian J. Walder, Spiridon I. Penev, Richard NockICML 2023 · 1 citation
- How to Boost Any Loss FunctionRichard Nock, Yishay MansourNeurIPS 2024
Related papers
- Learning to Approximate a Bregman DivergenceAli Siahkamari, Xide Xia, Venkatesh Saligrama, David A. Castañón et al.NeurIPS 2020 · 19 citations
- Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared LossAbhijeet Mulgund, Chirag PabbarajuICML 2025
- Omnipredicting Single-Index Models with Multi-index ModelsLunjia Hu, Kevin Tian, Chutong YangSTOC 2025 · 1 citation
- Binary Losses for Density Ratio EstimationWerner ZellingerICLR 2025
- On the sample complexity of semi-supervised multi-objective learningTobias Wegel, Geelon So, Junhyung Park, Fanny YangNeurIPS 2025 · 3 citations
