Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
Justin Chih-Yao Chen, Sukwon Yun, Elias Stengel-Eskin, Tianlong Chen, Mohit Bansal
摘要
Combining existing pre-trained LLMs is a promising approach for diverse reasoning tasks. However, task-level expert selection is often too coarsegrained, since different instances may require different expertise. To address this, we propose SKILL-MOE, a symbolic, skill-based, and gradient-free Mixture-of-Experts framework for instance-level expert selection. SKILL-MOE infers skills (e.g., algebra in mathematics) from each query, selects experts based on skill relevance, and lets each expert generate its own reasoning. The resulting k outputs are then synthesized by an aggregator chosen for its ability to integrate diverse responses. While instance-level selection substantially improves performance, naively implementing it incurs heavy overhead from repeated model loading and offloading. We address this with a batch inference strategy that groups instances by assigned experts, allowing each model to be loaded only once. As a result, SKILL-MOE integrates 16 expert models on a single GPU with runtime comparable to prior multi-agent baselines using 4 GPUs. Across diverse benchmarks (MMLU-Pro, GPQA, AIME, and MedMCQA), SKILL-MOE achieves an average absolute improvement of 8.15% over the best baseline. It also generalizes well to unseen tasks and outperforms discussion-based methods without requiring expensive multi-round interactions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficient Bilevel Optimization for CKA-Guided MoE UpcyclingZhiyuan Yu, Enneng Yang, Hao Jiang, Guojie Zhu 等ICML 2026
- The Avengers: A Routing Recipe for Collective Intelligence in Language ModelsYiqun Zhang, Hao Li, Chenxu Wang, Linyao Chen 等AAAI 2026
- TiME: Test-Time Mixture-of-Experts Routing via Asymmetric CO-Optimal Transport for Continual Test-Time AdaptationTianlun Liu, Zhiliang Tian, Zhen Huang, Tianle Liu 等ICML 2026
- Causal Dependency-Aware Unsupervised Routing for Large Reasoning ModelsJiacheng Liu, Hao Liu, Xiaofeng Hou, Wei Xue 等ICML 2026
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
相关 Paper
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang 等ICLR 2025
- DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-ExpertsJiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu 等ICML 2026 · 被引用 2 次
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding 等CVPR 2026 · 被引用 15 次
- S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous ReasoningJiangwen Dong, Zehui Lin, Wanyu Lin, Mingjin ZhangAAAI 2026 · 被引用 4 次
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng 等SC 2025 · 被引用 3 次
