Two-Stage Learning to Defer with Multiple Experts
Anqi Mao, Christopher Mohri, Mehryar Mohri, Yutao Zhong
Abstract
We study a two-stage scenario for learning to defer with multiple experts, which is crucial in practice for many applications. In this scenario, a predictor is derived in a first stage by training with a common loss function such as cross-entropy. In the second stage, a deferral function is learned to assign the most suitable expert to each input. We design a new family of surrogate loss functions for this scenario both in the score-based and the predictor-rejector settings and prove that they are supported by H-consistency bounds, which implies their Bayes-consistency. Moreover, we show that, for a constant cost function, our two-stage surrogate losses are realizable H-consistent. While the main focus of this work is a theoretical analysis, we also report the results of several experiments on CIFAR-10 and SVHN datasets. L hp exp 1 hp(x)=y ∑ ne i=1 e h(x,n+i)-maxy∈Y h(x,y) + ∑ ne j=1 c j (x, y) ∑ ne i=1,i≠j e h(x,n+i)-h(x,n+j) + e maxy∈Y h(x,y)-h(x,n+j) log -1 hp(x)=y log e max y∈Y h(x,y) e max y∈Y h(x,y) +∑ ne i=1 e h(x,n+i) -∑ ne j=1 c j (x, y) log e h(x,n+j) e max y∈Y h(x,y) +∑ ne i=1 e h(x,n+i) gce 1 hp(x)=y 1 α 1 -e max y∈Y h(x,y) e max y∈Y h(x,y) +∑ ne i=1 e h(x,n+i) α + ∑ ne j=1 c j (x, y) 1 α 1 -e h(x,n+j) e max y∈Y h(x,y) +∑ ne i=1 e h(x,n+i) α mae 1 hp(x)=y 1 -e max y∈Y h(x,y) e max y∈Y h(x,y) +∑ ne i=1 e h(x,n+i) + ∑ ne j=1 c j (x, y) 1 -e h(x,n+j) e max y∈Y h(x,y) +∑ ne i=1 e h( x,n+i)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eca10717-2a06-404d-8306-e0b8f1516e5fCited by top-tier papers32
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- Realizable H-Consistent and Bayes-Consistent Loss Functions for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2024 · 37 citations
- Structured Prediction with Stronger Consistency GuaranteesAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 37 citations
- H-Consistency Bounds: Characterization and ExtensionsAnqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 34 citations
Builds on22
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Is the Most Accurate AI the Best Teammate? Optimizing AI for TeamworkGagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz et al.AAAI 2021 · 185 citations
- Differentiable Learning Under TriageNastaran Okati, Abir De, Manuel Gomez-RodriguezNeurIPS 2021 · 99 citations
- Combining Human Predictions with Model Probabilities via Confusion Matrices and CalibrationGavin Kerrigan, Padhraic Smyth, Mark SteyversNeurIPS 2021 · 79 citations
Related papers
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to DeferAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Regression with Multi-Expert DeferralAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2024 · 31 citations
- A Two-Stage Learning-to-Defer Approach for Multi-Task LearningYannis Montreuil, Yeo Shu Heng, Axel Carlier, Lai Xing Ng et al.ICML 2025
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k ExpertsYannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang OoiICLR 2026 · 7 citations
- Exploiting Human-AI Dependence for Learning to DeferZixi Wei, Yuzhou Cao, Lei FengICML 2024 · 15 citations
