Bayes Consistency vs. H-Consistency: The Interplay between Surrogate Loss Functions and the Scoring Function Class
Mingyuan Zhang, Shivani Agarwal
Abstract
A fundamental question in multiclass classification concerns understanding the consistency properties of surrogate risk minimization algorithms, which minimize a (often convex) surrogate to the multiclass 0–1 loss. In particular, the framework of calibrated surrogates has played an important role in analyzing Bayes consistency of such algorithms, i.e. in studying convergence to a Bayes optimal classifier (Zhang, 2004; Tewari and Bartlett, 2007). However, follow-up work has suggested this framework can be of limited value when studying H-consistency; in particular, concerns have been raised that even when the data comes from an underlying linear model, minimizing certain convex calibrated surrogates over linear scoring functions fails to recover the true model (Long and Servedio, 2013). In this paper, we investigate this apparent conundrum. We find that while some calibrated surrogates can indeed fail to provide H-consistency when minimized over a naturallooking but naïvely chosen scoring function class F, the situation can potentially be remedied by minimizing them over a more carefully chosen class of scoring functions F. In particular, for the popular one-vs-all hinge and logistic surrogates, both of which are calibrated (and therefore provide Bayes consistency) under realizable models, but were previously shown to pose problems for realizable H-consistency, we derive a form of scoring function class F that enables Hconsistency. When H is the class of linear models, the class F consists of certain piecewise linear scoring functions that are characterized by the same number of parameters as in the linear case, and minimization over which can be performed using an adaptation of the min-pooling idea from neural network training. Our experiments confirm that the one-vs-all surrogates, when trained over this class of nonlinear scoring functions F, yield better linear multiclass classifiers than when trained over standard linear scoring functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
- H-Consistency Bounds for Surrogate Loss MinimizersPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2022 · 50 citations
- Multi-Class -Consistency BoundsPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2022 · 48 citations
- Realizable H-Consistent and Bayes-Consistent Loss Functions for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 37 citations
Related papers
- Calibration and Consistency of Adversarial Surrogate LossesPranjal Awasthi, Natalie Frank, Anqi Mao, Mehryar Mohri et al.NeurIPS 2021 · 59 citations
- On the consistency of top-k surrogate lossesForest Yang, Sanmi KoyejoICML 2020 · 54 citations
- Multi-Label Learning with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 30 citations
- Towards Consistency in Adversarial ClassificationLaurent Meunier, Raphael Ettedgui, Rafael Pinot, Yann Chevaleyre et al.NeurIPS 2022 · 12 citations
- A Universal Growth Rate for Learning with Smooth Surrogate LossesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 27 citations
