Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
Mengru Wang, Xingyu Chen, Yue Wang, Zhiwei He, Jiahao Xu, Tian Liang, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang, Ruotian Ma, Haitao Mi, Ningyu Zhang
摘要
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning performance without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed ''cognitive experts'' that orchestrate meta-level reasoning operations characterized by tokens like ''''. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu 等ICLR 2026 · 被引用 34 次
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt 等ICLR 2026 · 被引用 28 次
- Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoEYuanteng Chen, Peisong Wang, Nanxin Zeng, Yuantian Shao 等ICML 2026 · 被引用 3 次
- Rewiring Experts on the Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert ModelsGuinan Su, Yanwu Yang, Li Shen, Lu Yin 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper8
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni 等ICML 2024 · 被引用 183 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li 等EMNLP 2025 · 被引用 33 次
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt 等ICLR 2026 · 被引用 28 次
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired SketchingSimon A. Aytes, Jinheon Baek, Sung Ju HwangEMNLP 2025 · 被引用 3 次
相关 Paper
- Mixing Inference-time Experts for Enhancing LLM ReasoningSoumya Sanyal, Tianyi Xiao, Xiang RenEMNLP 2025
- MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language ModelsNurbek Tastan, Stefanos Laskaridis, Karthik Nandakumar, Samuel HorváthICML 2026 · 被引用 3 次
- How Far Are We from Optimal Reasoning Efficiency?Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang 等NeurIPS 2025 · 被引用 12 次
- Deft Scheduling of Dynamic Cloud Workflows with Varying Deadlines via Mixture-of-ExpertsYa Shen, Gang Chen, Hui Ma, Mengjie ZhangICLR 2026
- AdaMix: Adaptive Mixing for Short and Long Reasoning AdaptersHao Luo, Xiao Yan, Xinyan Li, Qiming Zeng 等ACL 2026
