Learning to Route Among Specialized Experts for Zero-Shot Generalization
Mohammed Muqeeth, Haokun Liu, Yufan Liu, Colin Raffel
Abstract
Recently, there has been a widespread proliferation of"expert"language models that are specialized to a specific task or domain through parameter-efficient fine-tuning. How can we recycle large collections of expert language models to improve zero-shot generalization to unseen tasks? In this work, we propose Post-Hoc Adaptive Tokenwise Gating Over an Ocean of Specialized Experts (PHATGOOSE), which learns to route among specialized modules that were produced through parameter-efficient fine-tuning. Unlike past methods that learn to route among specialized models, PHATGOOSE explores the possibility that zero-shot generalization will be improved if different experts can be adaptively chosen for each token and at each layer in the model. Crucially, our method is post-hoc - it does not require simultaneous access to the datasets used to create the specialized models and only requires a modest amount of additional compute after each expert model is trained. In experiments covering a range of specialized model collections and zero-shot generalization benchmarks, we find that PHATGOOSE outperforms past methods for post-hoc routing and, in some cases, outperforms explicit multitask training (which requires simultaneous data access). To better understand the routing strategy learned by PHATGOOSE, we perform qualitative experiments to validate that PHATGOOSE's performance stems from its ability to make adaptive per-token and per-module expert choices. We release all of our code to support future work on improving zero-shot generalization by recycling specialized experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e833ac0-6bc9-4ece-bc1f-e23a92d116f8Cited by top-tier papers27
- Towards Modular LLMs by Building and Reusing a Library of LoRAsOleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin et al.ICML 2024 · 70 citations
- Smoothie: Label Free Language Model RoutingNeel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare et al.NeurIPS 2024 · 44 citations
- When One LLM Drools, Multi-LLM Collaboration RulesShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang et al.ACL 2026 · 27 citations
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention RoutingWenbing Li, Zikai Song, Hang Zhou, Junqing Yu et al.ICLR 2026 · 20 citations
- Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time ComputeJianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi et al.AAAI 2026 · 18 citations
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task LearningMd Kowsher, Haris Mansoor, Nusrat Prottasha, Ozlem Garibay et al.ICML 2026
- Multi-Head Adapter Routing for Cross-Task GeneralizationLucas Page-Caccia, Edoardo Maria Ponti, Zhan Su, Matheus Pereira et al.NeurIPS 2023 · 38 citations
- Split-Merge: Scalable and Memory-Efficient Merging of Expert LLMsSruthi Gorantla, Aditya Rawal, Devamanyu Hazarika, Kaixiang Lin et al.EMNLP 2025
- Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsZihan Wang, Deli Chen, Damai Dai, Runxin Xu et al.EMNLP 2024 · 2 citations
- Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-TrainingReza Shirkavand, Peiran Yu, Qi He, Heng HuangNeurIPS 2025 · 6 citations
