Toward Calibrated Mixture-of-Experts Under Distribution Shift
Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa, Anqi Liu
摘要
Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
相关 Paper
- When Model Merging Breaks Routing: Training-Free Calibration for MoECanbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang 等ICML 2026
- Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts TransformersAlbus Li, Matthew WickerICML 2026 · 被引用 2 次
- -Balancing for Mixture-of-Experts TrainingLizhang Chen, Jonathan Li, Qi Wang, Runlong Liao 等ICML 2026
- Improving Calibration through the Relationship with Adversarial RobustnessYao Qin, Xuezhi Wang, Alex Beutel, Ed H. ChiNeurIPS 2021 · 被引用 32 次
- Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQAMahsa Mozaffari, Hitesh Sapkota, Yu Kong, Xumin Liu 等ICML 2026
