Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured Prediction
Mehryar Mohri, Yutao Zhong
摘要
A fundamental dichotomy in the theory of classification sets smoothness against statistical efficiency: smooth surrogate losses such as the logistic loss enable fast optimization but yield slow square-root -consistency bounds, while piecewise-linear losses like the Hinge loss achieve optimal linear -consistency rates but are non-differentiable. We introduce Linear-Core (LC) Surrogates, the first family of explicit convex loss functions that provably resolve this tension. By stitching a linear core to a smooth tail, we construct surrogates that are differentiable everywhere (, and even under mild conditions) while retaining strict linear -consistency bounds, the strongest known form of consistency guarantee. We establish these linear bounds across three increasingly complex settings: binary classification, multi-class classification, and structured prediction. To our knowledge, this is the first explicit construction to simultaneously achieve smoothness and linear -consistency in any of these settings. Beyond their theoretical appeal, Linear-Core Surrogates offer practical advantages. In multi-class classification, their constant gradient profile near the decision boundary provides natural robustness to instance-dependent label noise, outperforming Cross-Entropy by 2.6% on corrupted CIFAR-10. In structured prediction, their smoothness enables an unbiased stochastic gradient estimator that bypasses the per-step complexity of exact inference, yielding a 23 speedup over Structured SVMs on large-vocabulary sequence tagging tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k ExpertsYannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang OoiICLR 2026 · 被引用 7 次
- A Theoretical Framework for Modular Learning of Robust Generative ModelsCorinna Cortes, Mehryar Mohri, Yutao ZhongICML 2026 · 被引用 7 次
- Optimized Deferral for Imbalanced SettingsCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2026 · 被引用 7 次
- Mind the Gap: Structure-Aware Consistency in Preference LearningMehryar Mohri, Yutao ZhongICML 2026
它引用的顶会 Paper28
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 被引用 790 次
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang 等NeurIPS 2020 · 被引用 329 次
- Learning with Bounded Instance and Label-dependent Label NoiseJiacheng Cheng, Tongliang Liu, Kotagiri Ramamohanarao, Dacheng TaoICML 2020 · 被引用 162 次
- Confidence Scores Make Instance-dependent Label-noise Learning PossibleAntonin Berthon, Bo Han, Gang Niu, Tongliang Liu 等ICML 2021 · 被引用 126 次
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 被引用 98 次
相关 Paper
- Structured Prediction with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 被引用 37 次
- Multi-Class -Consistency BoundsPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2022 · 被引用 48 次
- Multi-Label Learning with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 被引用 30 次
- H-Consistency Bounds for Surrogate Loss MinimizersPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2022 · 被引用 50 次
- A Universal Growth Rate for Learning with Smooth Surrogate LossesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 被引用 27 次
