SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input Routing
Jaehan Kim, Minkyoo Song, Seungwon Shin, Sooel Son
Abstract
Recent large language models (LLMs) have increasingly adopted the Mixture-of-Experts (MoE) architecture for efficiency. MoE-based LLMs heavily depend on a superficial safety mechanism in which harmful inputs are routed safety-critical experts. However, our analysis reveals that routing decisions for harmful inputs drift significantly after fine-tuning, exposing a critical vulnerability to harmful fine-tuning (HFT) attacks. Existing defenses, primarily designed for monolithic LLMs, are less effective for MoE LLMs as they fail to prevent drift in harmful input routing. To address this limitation, we propose SafeMoE, a safe fine-tuning method tailored to MoE LLMs. SafeMoE directly mitigates routing drift by penalizing the gap between the routing weights of a fine-tuned model and those of the initial safety-aligned model, thereby preserving the safety-aligned routing of harmful inputs to safety-critical experts. Experiments on open-source MoE LLMs ranging from 7B to 141B parameters demonstrate that SafeMoE effectively mitigates HFT attacks, reducing the harmfulness score of OLMoE from 62.0 to 5.0, for example, while maintaining task utility within 1% degradation and incurring only 2% overhead. It significantly outperforms state-of-the-art defense methods for safeguarding LLM fine-tuning and remains effective in recent large-scale MoE LLMs such as gpt-oss and Llama 4. Our implementation is available at https://github.com/jaehanwork/SafeMoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on29
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsYukun Jiang, Hai Huang, Mingjie Li, Yage Zhang et al.ICML 2026 · 9 citations
- MESA: Improving MoE Safety Alignment via Decentralized ExpertiseYitong Sun, Yao Huang, Teng Li, Ranjie Duan et al.ICML 2026
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert IdentificationZhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu et al.NeurIPS 2025 · 22 citations
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Stjepan Picek et al.USENIX Security 2026 · 14 citations
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts TransformersXin Zhao, Xiaojun Chen, Bingshan Liu, Haoyu Gao et al.NeurIPS 2025 · 3 citations
