Toward Calibrated Mixture-of-Experts Under Distribution Shift
Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa, Anqi Liu
Abstract
Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
Related papers
- When Model Merging Breaks Routing: Training-Free Calibration for MoECanbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang et al.ICML 2026
- Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts TransformersAlbus Li, Matthew WickerICML 2026 · 2 citations
- -Balancing for Mixture-of-Experts TrainingLizhang Chen, Jonathan Li, Qi Wang, Runlong Liao et al.ICML 2026
- Improving Calibration through the Relationship with Adversarial RobustnessYao Qin, Xuezhi Wang, Alex Beutel, Ed H. ChiNeurIPS 2021 · 32 citations
- Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQAMahsa Mozaffari, Hitesh Sapkota, Yu Kong, Xumin Liu et al.ICML 2026
