On the Adversarial Robustness of Mixture of Experts
Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme, Pranjal Awasthi, Srinadh Bhojanapalli
摘要
Adversarial robustness is a key desirable property of neural networks. It has been empirically shown to be affected by their sizes, with larger networks being typically more robust. Recently, Bubeck and Sellke proved a lower bound on the Lipschitz constant of functions that fit the training data in terms of their number of parameters. This raises an interesting open question, do -- and can -- functions with more parameters, but not necessarily more computational cost, have better robustness? We study this question for sparse Mixture of Expert models (MoEs), that make it possible to scale up the model size for a roughly constant computational cost. We theoretically show that under certain conditions on the routing and the structure of the data, MoEs can have significantly smaller Lipschitz constants than their dense counterparts. The robustness of MoEs can suffer when the highest weighted experts for an input implement sufficiently different functions. We next empirically evaluate the robustness of MoEs on ImageNet using adversarial attacks and show they are indeed more robust than dense models with the same computational cost. We make key observations showing the robustness of MoEs to the choice of experts, highlighting the redundancy of experts in models trained in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Robust Mixture-of-Expert Training for Convolutional Neural NetworksYihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang 等ICCV 2023 · 被引用 43 次
- GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned ExpertsShirley Wu, Kaidi Cao, Bruno Ribeiro, James Y. Zou 等NeurIPS 2024 · 被引用 27 次
- MomentumSMoE: Integrating Momentum into Sparse Mixture of ExpertsRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 被引用 10 次
- SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?Haomin Zhuang, Yihua Zhang, Kehan Guo, Jinghan Jia 等ACL 2025 · 被引用 10 次
- Sparsity-Constrained Optimal TransportTianlin Liu, Joan Puigcerver, Mathieu BlondelICLR 2023 · 被引用 3 次
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
相关 Paper
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 被引用 1 次
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 被引用 27 次
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang 等ICML 2026
- A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-ExpertsMohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang 等ICML 2024 · 被引用 18 次
