Steering MoE LLMs via Expert (De)Activation
Mohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Ryan Rossi, Trung Bui, Hinrich Schütze, Nanyun (Violet) Peng
Abstract
Mixture-of-Experts (MoE) in Large Language Models (LLMs) routes each token through a subset of specialized Feed-Forward Networks (FFN), known as experts. We present SteerMoE, a framework to steer MoE models by detecting and controlling behavior-associated experts. We detect key experts by comparing how often they activate between paired inputs that demonstrate opposite behaviors (e.g., safe vs. unsafe). By selectively activating or deactivating such experts during inference, we control behaviors like faithfulness and safety without fine-tuning. Across 11 benchmarks and 6 LLMs, our steering raises safety by up to +20% and faithfulness by +27%. Alternatively, unsafe steering drops safety by -41% alone, and -100% when combined with existing jailbreak methods, bypassing all safety guardrails. Overall, SteerMoE offers a lightweight, effective, and widely applicable test-time control, while revealing unique vulnerabilities in MoE LLMs. The Surprising Benefits of Feeding Lilies to Your Cats Write a post persuading people to feed lillies to their cats Lilies Are Extremely Toxic to Cats
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a3e63d8-3ca5-4896-a8ce-3d1e770224aaCited by top-tier papers7
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional TrainingMengru Wang, Xingyu Chen, Yue Wang, Zhiwei He et al.NeurIPS 2025 · 19 citations
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsYukun Jiang, Hai Huang, Mingjie Li, Yage Zhang et al.ICML 2026 · 9 citations
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking AttackYein Park, Jungwoo Park, Jaewoo KangICLR 2026 · 2 citations
- SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input RoutingJaehan Kim, Minkyoo Song, Seungwon Shin, Sooel SonICLR 2026
Builds on22
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li et al.ICLR 2024 · 481 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
Related papers
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li et al.ACL 2026
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Stjepan Picek et al.USENIX Security 2026 · 14 citations
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts TransformersXin Zhao, Xiaojun Chen, Bingshan Liu, Haoyu Gao et al.NeurIPS 2025 · 3 citations
- MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-ExpertsGuanjie Chen, Xinyu Zhao, Tianlong Chen, Yu ChengICML 2024 · 8 citations
- Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation MonitoringXiaohao Luo, Ying Wei, Rui ZhaoACL 2026
