Is the acquisition worth the cost? Surrogate losses for Consistent Two-stage Classifiers
Florence Regol, Joseph Cotnareanu, Theodore Glavas, Mark Coates
摘要
Recent years have witnessed the emergence of a spectrum of foundation models, covering a broad range of capabilities and costs. Often, we effectively use foundation models as feature generators and train classifiers that use the outputs of these models to make decisions. In this paper, we consider an increasingly relevant setting where we have two classifier stages. The first stage has access to features x and has the option to make a classification decision or defer, while incurring a cost, to a second classifier that has access to features x and z . This is similar to the “learning to defer” setting, with the important difference that we train both classifiers jointly, and the second classifier has access to more information. The natural loss for this setting is an ℓ 01 c loss, where a penalty is paid for incorrect classification, as in ℓ 01 , but an additional penalty c is paid for consulting the second classifier. The ℓ 01 c loss is unwieldy for training. Our primary contribution in this paper is the derivation of a hinge-based surrogate loss ℓ chinge that is much more amenable to training but also satisfies the property that ℓ chinge -consistency implies ℓ 01 c -consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou 等ICML 2023 · 被引用 318 次
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 被引用 267 次
相关 Paper
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 被引用 98 次
- Exploiting Human-AI Dependence for Learning to DeferZixi Wei, Yuzhou Cao, Lei FengICML 2024 · 被引用 15 次
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Learning to Help in Multi-Class SettingsYu Wu, Yansong Li, Zeyu Dong, Nitya Sathyavageeswaran 等ICLR 2025
- Post-hoc estimators for learning to defer to an expertHarikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat 等NeurIPS 2022 · 被引用 66 次
