Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured Prediction
Mehryar Mohri, Yutao Zhong
Abstract
A fundamental dichotomy in the theory of classification sets smoothness against statistical efficiency: smooth surrogate losses such as the logistic loss enable fast optimization but yield slow square-root -consistency bounds, while piecewise-linear losses like the Hinge loss achieve optimal linear -consistency rates but are non-differentiable. We introduce Linear-Core (LC) Surrogates, the first family of explicit convex loss functions that provably resolve this tension. By stitching a linear core to a smooth tail, we construct surrogates that are differentiable everywhere (, and even under mild conditions) while retaining strict linear -consistency bounds, the strongest known form of consistency guarantee. We establish these linear bounds across three increasingly complex settings: binary classification, multi-class classification, and structured prediction. To our knowledge, this is the first explicit construction to simultaneously achieve smoothness and linear -consistency in any of these settings. Beyond their theoretical appeal, Linear-Core Surrogates offer practical advantages. In multi-class classification, their constant gradient profile near the decision boundary provides natural robustness to instance-dependent label noise, outperforming Cross-Entropy by 2.6% on corrupted CIFAR-10. In structured prediction, their smoothness enables an unbiased stochastic gradient estimator that bypasses the per-step complexity of exact inference, yielding a 23 speedup over Structured SVMs on large-vocabulary sequence tagging tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k ExpertsYannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang OoiICLR 2026 · 7 citations
- A Theoretical Framework for Modular Learning of Robust Generative ModelsCorinna Cortes, Mehryar Mohri, Yutao ZhongICML 2026 · 7 citations
- Optimized Deferral for Imbalanced SettingsCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2026 · 7 citations
- Mind the Gap: Structure-Aware Consistency in Preference LearningMehryar Mohri, Yutao ZhongICML 2026
Builds on28
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang et al.NeurIPS 2020 · 329 citations
- Learning with Bounded Instance and Label-dependent Label NoiseJiacheng Cheng, Tongliang Liu, Kotagiri Ramamohanarao, Dacheng TaoICML 2020 · 162 citations
- Confidence Scores Make Instance-dependent Label-noise Learning PossibleAntonin Berthon, Bo Han, Gang Niu, Tongliang Liu et al.ICML 2021 · 126 citations
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
Related papers
- Structured Prediction with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 37 citations
- Multi-Class -Consistency BoundsPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2022 · 48 citations
- Multi-Label Learning with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 30 citations
- H-Consistency Bounds for Surrogate Loss MinimizersPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2022 · 50 citations
- A Universal Growth Rate for Learning with Smooth Surrogate LossesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 27 citations
