Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach
Xu Zhang, Kaidi Xu, Ziqing Hu, Ren Wang
Abstract
Mixture of Experts (MoE) have shown remarkable success in leveraging specialized expert networks for complex machine learning tasks. However, their susceptibility to adversarial attacks presents a critical challenge for deployment in robust applications. This paper addresses the critical question of how to incorporate robustness into MoEs while maintaining high natural accuracy. We begin by analyzing the vulnerability of MoE components, finding that expert networks are notably more susceptible to adversarial attacks than the router. Based on this insight, we propose a targeted robust training technique that integrates a novel loss function to enhance the adversarial robustness of MoE, requiring only the robustification of one additional expert without compromising training or inference efficiency. Building on this, we introduce a dual-model strategy that linearly combines a standard MoE model with our robustified MoE model using a smoothing parameter. This approach allows for flexible control over the robustness-accuracy trade-off. We further provide theoretical foundations by deriving certified robustness bounds for both the single MoE and the dual-model. To push the boundaries of robustness and accuracy, we propose a novel joint training strategy JTDMoE for the dual-model. This joint training enhances both robustness and accuracy beyond what is achievable with separate models. Experimental results on CIFAR-10 and TinyImageNet datasets using ResNet18 and Vision Transformer (ViT) architectures demonstrate the effectiveness of our proposed methods. The code is publicly available at https://github.com/TIML-Group/ Robust-MoE-Dual-Model .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Exposing and Defending the Achilles' Heel of Video Mixture-of-ExpertsSongping Wang, Qinglong Liu, Yueming Lyu, Ning Li et al.ICLR 2026 · 3 citations
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 1 citation
- MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary AttacksHyo Seo Kim, Gang Luo, Can Chen, Binghui Wang et al.ICML 2026
Builds on12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
Related papers
- Robust Mixture-of-Expert Training for Convolutional Neural NetworksYihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang et al.ICCV 2023 · 43 citations
- ReMoE: Region-Mixture Experts for Adversarially-Robust Vision TransformersQinghao Zhong, Bingzhi Chen, Yishu Liu, Minhua Lu et al.CVPR 2026
- On the Adversarial Robustness of Mixture of ExpertsJoan Puigcerver, Rodolphe Jenatton, Carlos Riquelme, Pranjal Awasthi et al.NeurIPS 2022 · 33 citations
- Revisiting adapters with adversarial trainingSylvestre-Alvise Rebuffi, Francesco Croce, Sven GowalICLR 2023
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
