Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
Pierre Ablin, Angelos Katharopoulos, Skyler Seto, David Grangier
Abstract
Machine learning models are routinely trained on a mixture of different data domains. Different domain weights yield very different downstream performances. We propose the Soup-of-Experts, a novel architecture that can instantiate a model at test time for any domain weights with minimal computational cost and without re-training the model. Our architecture consists of a bank of expert parameters, which are linearly combined to instantiate one model. We learn the linear combination coefficients as a function of the input domain weights. To train this architecture, we sample random domain weights, instantiate the corresponding model, and backprop through one batch of data sampled with these domain weights. We demonstrate how our approach obtains small specialized models on several language modeling tasks quickly. Soup-of-Experts are particularly appealing when one needs to ship many different specialist models quickly under a model size constraint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b83bc8bb-9a23-461f-97ee-ac4d5796c621Cited by top-tier papers2
- Self-Soupervision: Cooking Model Soups without LabelsAnthony Fuller, James Green, Evan ShelhamerICML 2026
- One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative PerceptionYang Li, Weize Li, Quan Yuan, Shao Congzhang et al.ICML 2026
Builds on15
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 236 citations
Related papers
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural NetworksHakbin Kim, Dong-Wan ChoiSIGMOD 2021 · 1 citation
- A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She et al.ACL 2026
- Generalizability of Mixture of Domain-Specific Adapters from the Lens of Signed Weight Directions and its Application to Effective Model PruningTuc Nguyen, Thai LeACL 2024
- No Need to Talk: Asynchronous Mixture of Language ModelsAnastasiia Filippova, Angelos Katharopoulos, David Grangier, Ronan CollobertICLR 2025
- XPERT: Expert Knowledge Transfer for Effective Training of Language ModelsChang Liu, boyu shi, Xu Yang, Xin GengICML 2026 · 2 citations
