All your loss are belong to Bayes
Christian J. Walder, Richard Nock
Abstract
Loss functions are a cornerstone of machine learning and the starting point of most algorithms. Statistics and Bayesian decision theory have contributed, via properness, to elicit over the past decades a wide set of admissible losses in supervised learning, to which most popular choices belong (logistic, square, Matsushita, etc.). Rather than making a potentially biased ad hoc choice of the loss, there has recently been a boost in efforts to fit the loss to the domain at hand while training the model itself. The key approaches fit a canonical link, a function which monotonically relates the closed unit interval to R and can provide a proper loss via integration. In this paper, we rely on a broader view of proper composite losses and a recent construct from information geometry, source functions, whose fitting alleviates constraints faced by canonical links. We introduce a trick on squared Gaussian Processes to obtain a random process whose paths are compliant source functions with many desirable properties in the context of link estimation. Experimental results demonstrate substantial improvements over the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- LegendreTron: Uprising Proper Multiclass Loss LearningKevin H. Lam, Christian J. Walder, Spiridon I. Penev, Richard NockICML 2023 · 1 citation
- DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and Its Loss' Convexity is Dispensable)Wenxuan Zhou, Shujian Zhang, brice magdalou, John Lambert et al.ICML 2026
Builds on1
Related papers
- Learning with Fitzpatrick LossesSeta Rakotomandimby, Jean-Philippe Chancelier, Michel De Lara, Mathieu BlondelNeurIPS 2024 · 6 citations
- Empirical Gaussian ProcessesJihao Andreas Lin, Sebastian Ament, Louis Tiao, David Eriksson et al.ICML 2026
- Binary Losses for Density Ratio EstimationWerner ZellingerICLR 2025
- Fast Bayesian Inference for Gaussian Cox Processes via Path Integral FormulationHideaki KimNeurIPS 2021 · 8 citations
- How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age EstimationAli Akbari, Muhammad Awais, Manijeh Bashar, Josef KittlerICML 2021 · 46 citations
