DEM: Distribution Edited Model for Training with Mixed Data Distributions
Dhananjay Ram, Aditya Rawal, Momchil Hardalov, Nikolaos Pappas, Sheng Zha
Abstract
Training with mixed data distributions is a common and important part of creating multi-task and instruction-following models.The diversity of the data distributions and cost of joint training makes the optimization procedure extremely challenging.Data mixing methods partially address this problem, albeit having a sub-optimal performance across data sources and require multiple expensive training runs.In this paper, we propose a simple and efficient alternative for better optimization of the data sources by combining models individually trained on each data source with the base model using basic element-wise vector operations.The resulting model, namely Distribution Edited Model (DEM), is 11 cheaper than standard data mixing and outperforms strong baselines on a variety of benchmarks, yielding upto 6.2% improvement on MMLU, 11.5% on BBH, 16.1% on DROP, 6% on MathQA, and 9.3% on HELM with models of size 3B to 13B.Notably, DEM does not require full re-training when modifying a single data-source, thus making it very flexible and scalable for training with diverse data sources.The code is available at https://github.com/amazon-science/demdistribution-edited-model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb4b4d0b-7dab-49d5-b6a5-6b9f817741e4Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye et al.ICML 2026 · 3 citations
- Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task AdaptationChen Dun, Mirian del Carmen Hipolito Garcia, Guoqing Zheng, Ahmed Hassan Awadallah et al.AAAI 2025 · 7 citations
- Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without TrainingMozhi Zhang, Howe Tissue, Lu Wang, Xipeng QiuICML 2025
- MixMin: Finding Data Mixtures via Convex MinimizationAnvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush et al.ICML 2025
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur et al.SIGMOD 2026 · 5 citations
