DEM: Distribution Edited Model for Training with Mixed Data Distributions
Dhananjay Ram, Aditya Rawal, Momchil Hardalov, Nikolaos Pappas, Sheng Zha
摘要
Training with mixed data distributions is a common and important part of creating multi-task and instruction-following models.The diversity of the data distributions and cost of joint training makes the optimization procedure extremely challenging.Data mixing methods partially address this problem, albeit having a sub-optimal performance across data sources and require multiple expensive training runs.In this paper, we propose a simple and efficient alternative for better optimization of the data sources by combining models individually trained on each data source with the base model using basic element-wise vector operations.The resulting model, namely Distribution Edited Model (DEM), is 11 cheaper than standard data mixing and outperforms strong baselines on a variety of benchmarks, yielding upto 6.2% improvement on MMLU, 11.5% on BBH, 16.1% on DROP, 6% on MathQA, and 9.3% on HELM with models of size 3B to 13B.Notably, DEM does not require full re-training when modifying a single data-source, thus making it very flexible and scalable for training with diverse data sources.The code is available at https://github.com/amazon-science/demdistribution-edited-model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye 等ICML 2026 · 被引用 3 次
- Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task AdaptationChen Dun, Mirian del Carmen Hipolito Garcia, Guoqing Zheng, Ahmed Hassan Awadallah 等AAAI 2025 · 被引用 7 次
- Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without TrainingMozhi Zhang, Howe Tissue, Lu Wang, Xipeng QiuICML 2025
- MixMin: Finding Data Mixtures via Convex MinimizationAnvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush 等ICML 2025
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur 等SIGMOD 2026 · 被引用 5 次
