Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
Justin Chih-Yao Chen, Sukwon Yun, Elias Stengel-Eskin, Tianlong Chen, Mohit Bansal
Abstract
Combining existing pre-trained LLMs is a promising approach for diverse reasoning tasks. However, task-level expert selection is often too coarsegrained, since different instances may require different expertise. To address this, we propose SKILL-MOE, a symbolic, skill-based, and gradient-free Mixture-of-Experts framework for instance-level expert selection. SKILL-MOE infers skills (e.g., algebra in mathematics) from each query, selects experts based on skill relevance, and lets each expert generate its own reasoning. The resulting k outputs are then synthesized by an aggregator chosen for its ability to integrate diverse responses. While instance-level selection substantially improves performance, naively implementing it incurs heavy overhead from repeated model loading and offloading. We address this with a batch inference strategy that groups instances by assigned experts, allowing each model to be loaded only once. As a result, SKILL-MOE integrates 16 expert models on a single GPU with runtime comparable to prior multi-agent baselines using 4 GPUs. Across diverse benchmarks (MMLU-Pro, GPQA, AIME, and MedMCQA), SKILL-MOE achieves an average absolute improvement of 8.15% over the best baseline. It also generalizes well to unseen tasks and outperforms discussion-based methods without requiring expensive multi-round interactions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 052b5a07-7d26-487c-acb1-e441f9b8021eCited by top-tier papers4
- Efficient Bilevel Optimization for CKA-Guided MoE UpcyclingZhiyuan Yu, Enneng Yang, Hao Jiang, Guojie Zhu et al.ICML 2026
- The Avengers: A Routing Recipe for Collective Intelligence in Language ModelsYiqun Zhang, Hao Li, Chenxu Wang, Linyao Chen et al.AAAI 2026
- TiME: Test-Time Mixture-of-Experts Routing via Asymmetric CO-Optimal Transport for Continual Test-Time AdaptationTianlun Liu, Zhiliang Tian, Zhen Huang, Tianle Liu et al.ICML 2026
- Causal Dependency-Aware Unsupervised Routing for Large Reasoning ModelsJiacheng Liu, Hao Liu, Xiaofeng Hou, Wei Xue et al.ICML 2026
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang et al.ICLR 2025
- DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-ExpertsJiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu et al.ICML 2026 · 2 citations
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding et al.CVPR 2026 · 15 citations
- S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous ReasoningJiangwen Dong, Zehui Lin, Wanyu Lin, Mingjin ZhangAAAI 2026 · 4 citations
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng et al.SC 2025 · 3 citations
