No Need to Talk: Asynchronous Mixture of Language Models
Anastasiia Filippova, Angelos Katharopoulos, David Grangier, Ronan Collobert
摘要
We introduce SMALLTALK LM, an innovative method for training a mixture of language models in an almost asynchronous manner. Each model of the mixture specializes in distinct parts of the data distribution, without the need for high-bandwidth communication between the nodes training each model. At inference, a lightweight router directs a given sequence to a single expert, according to a short prefix. This inference scheme naturally uses a fraction of the parameters from the overall mixture model. Unlike prior works on asynchronous LLM training, our routing method does not rely on full corpus clustering or access to metadata, making it more suitable for real-world applications. Our experiments on language modeling demonstrate that SMALLTALK LM achieves significantly lower perplexity than dense model baselines for the same total training FLOPs and an almost identical inference cost. Finally, in our downstream evaluations we outperform the dense baseline on 75% of the tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Optimal Splitting of Language Models from Mixtures to Specialized DomainsSkyler Seto, Pierre Ablin, Anastasiia Filippova, Jiayuan Ye 等ICML 2026 · 被引用 2 次
- BTS: Harmonizing Specialized Experts into a Generalist LLMQizhen Zhang, Prajjwal Bhargava, Chloe Bi, Chris X. Cai 等EMNLP 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
相关 Paper
- Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu 等ACL 2026 · 被引用 6 次
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai 等NeurIPS 2022 · 被引用 223 次
- Language Model Networks: Supervision-Efficient Learning through Dense CommunicationShiguang Wu, Yaqing Wang, QUANMING YAOICML 2026 · 被引用 3 次
- DiSRouter: Distributed Self-Routing for LLM SelectionsHang Zheng, Hongshen Xu, Yongkai.lin, Shuai Fan 等ICLR 2026 · 被引用 6 次
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
