Is the acquisition worth the cost? Surrogate losses for Consistent Two-stage Classifiers
Florence Regol, Joseph Cotnareanu, Theodore Glavas, Mark Coates
Abstract
Recent years have witnessed the emergence of a spectrum of foundation models, covering a broad range of capabilities and costs. Often, we effectively use foundation models as feature generators and train classifiers that use the outputs of these models to make decisions. In this paper, we consider an increasingly relevant setting where we have two classifier stages. The first stage has access to features x and has the option to make a classification decision or defer, while incurring a cost, to a second classifier that has access to features x and z . This is similar to the “learning to defer” setting, with the important difference that we train both classifiers jointly, and the second classifier has access to more information. The natural loss for this setting is an ℓ 01 c loss, where a penalty is paid for incorrect classification, as in ℓ 01 , but an additional penalty c is paid for consulting the second classifier. The ℓ 01 c loss is unwieldy for training. Our primary contribution in this paper is the derivation of a hinge-based surrogate loss ℓ chinge that is much more amenable to training but also satisfies the property that ℓ chinge -consistency implies ℓ 01 c -consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bca550ce-8bfa-4a5f-b6c2-5c3e8979876fCited by top-tier papers1
Ask how each one uses itBuilds on16
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
Related papers
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
- Exploiting Human-AI Dependence for Learning to DeferZixi Wei, Yuzhou Cao, Lei FengICML 2024 · 15 citations
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Learning to Help in Multi-Class SettingsYu Wu, Yansong Li, Zeyu Dong, Nitya Sathyavageeswaran et al.ICLR 2025
- Post-hoc estimators for learning to defer to an expertHarikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat et al.NeurIPS 2022 · 66 citations
