BTS: Harmonizing Specialized Experts into a Generalist LLM
Qizhen Zhang, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob Nicolaus Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen, Emily Dinan, Suchin Gururangan, Mike Lewis
Abstract
We present BRANCH-TRAIN-STITCH (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable generalist model. Following Li et al. [2022], we start with a single seed LLM which is branched into domain-specific (e.g., coding or math) experts with continual pretraining. BTS combines experts using lightweight stitch layers, which are inserted between frozen experts and the seed LLM, and trained on a small datamix of the expert domains. Stitch layers enable the seed LLM to integrate representations from any number of experts during the forward pass, allowing it to generalize to new domains, despite remaining frozen. Because BTS does not alter the constituent LLMs, BTS provides a modular and flexible approach: experts can be easily removed or added with only a small amount of training. Compared to alternative model merging or upcycling approaches, BTS yields the best generalist performance on a variety of downstream tasks, while retaining the specialized capabilities of each of the experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 · 433 citations
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 140 citations
Related papers
- A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She et al.ACL 2026
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer ChunkingDengming Zhang, Xiaowen Ma, Zhenliang Ni, Zhenkai Wu et al.ICLR 2026 · 6 citations
- Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language ModelsLucas Bandarkar, Benjamin Muller, Pritish Yuvraj, Rui Hou et al.ICLR 2025
- XPERT: Expert Knowledge Transfer for Effective Training of Language ModelsChang Liu, boyu shi, Xu Yang, Xin GengICML 2026 · 2 citations
- Split-Merge: Scalable and Memory-Efficient Merging of Expert LLMsSruthi Gorantla, Aditya Rawal, Devamanyu Hazarika, Kaixiang Lin et al.EMNLP 2025
