Optimized Deferral for Imbalanced Settings
Corinna Cortes, Anqi Mao, Mehryar Mohri, Yutao Zhong
Abstract
Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. This approach, known as learning to defer , is essential in domains like natural language generation, medical diagnosis, and computer vision, where an effective deferral can reduce errors at low extra resource consumption. However, the two-stage learning to defer setting, which leverages existing predictors such as a collection of LLMs or other classifiers, often faces challenges due to an expert imbalance problem. This imbalance can lead to suboptimal performance, with deferral algorithms favoring the majority expert. We present a comprehensive study of two-stage learning to defer in expert imbalance settings. We cast the deferral loss optimization as a novel cost-sensitive learning problem over the input-expert domain. We derive new margin-based loss functions and guarantees tailored to this setting, and develop novel algorithms for cost-sensitive learning. Leveraging these results, we design principled deferral algorithms, MILD ( Margin-based Imbalanced Learning to Defer ), specifically suited for expert imbalance settings. Extensive experiments demonstrate the effectiveness of our approach, showing clear improvements over existing baselines on both image classification and real-world Large Language Model (LLM) routing tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6ba3731-debc-49a8-af24-bb31b887d09eCited by top-tier papers3
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k ExpertsYannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang OoiICLR 2026 · 7 citations
- Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured PredictionMehryar Mohri, Yutao ZhongICML 2026 · 7 citations
- Mind the Gap: Structure-Aware Consistency in Preference LearningMehryar Mohri, Yutao ZhongICML 2026
Builds on58
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
Related papers
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Post-hoc estimators for learning to defer to an expertHarikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat et al.NeurIPS 2022 · 66 citations
- Learning to Defer with Limited Expert PredictionsPatrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik et al.AAAI 2023 · 28 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
