Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
Mengru Wang, Xingyu Chen, Yue Wang, Zhiwei He, Jiahao Xu, Tian Liang, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang, Ruotian Ma, Haitao Mi, Ningyu Zhang
Abstract
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning performance without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed ''cognitive experts'' that orchestrate meta-level reasoning operations characterized by tokens like ''''. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d99b4bf8-4e46-4bad-a54d-8901e0763e92Cited by top-tier papers4
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoEYuanteng Chen, Peisong Wang, Nanxin Zeng, Yuantian Shao et al.ICML 2026 · 3 citations
- Rewiring Experts on the Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert ModelsGuinan Su, Yanwu Yang, Li Shen, Lu Yin et al.ICML 2026 · 3 citations
Builds on8
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni et al.ICML 2024 · 183 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li et al.EMNLP 2025 · 33 citations
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired SketchingSimon A. Aytes, Jinheon Baek, Sung Ju HwangEMNLP 2025 · 3 citations
Related papers
- Mixing Inference-time Experts for Enhancing LLM ReasoningSoumya Sanyal, Tianyi Xiao, Xiang RenEMNLP 2025
- MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language ModelsNurbek Tastan, Stefanos Laskaridis, Karthik Nandakumar, Samuel HorváthICML 2026 · 3 citations
- How Far Are We from Optimal Reasoning Efficiency?Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang et al.NeurIPS 2025 · 12 citations
- Deft Scheduling of Dynamic Cloud Workflows with Varying Deadlines via Mixture-of-ExpertsYa Shen, Gang Chen, Hui Ma, Mengjie ZhangICLR 2026
- AdaMix: Adaptive Mixing for Short and Long Reasoning AdaptersHao Luo, Xiao Yan, Xinyan Li, Qiming Zeng et al.ACL 2026
