MixMin: Finding Data Mixtures via Convex Minimization
Anvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush, Chris J. Maddison
Abstract
Modern machine learning pipelines are increasingly combining and mixing data from diverse and disparate sources, e.g., pre-training large language models. Yet, finding the optimal data mixture is a challenging and open problem. We formalize this data mixing problem as a bi-level objective: the best mixture is the one that would lead to the best model for a downstream objective. Unfortunately, this objective is generally intractable. In this paper, we make the observation that the bi-level data mixing objective becomes convex as our model class becomes larger. We develop and study a gradient-based approach for optimizing this convex objective, which we call MixMin, and test it on language modeling and chemistry tasks. MixMin was the only method that uniformly improved the data mixture in all our experiments. With MixMin, we improved the data mixture using less than 0.2% additional compute for a pythia-410M model trained on 8.2B tokens, resulting between 1-5% relative improvement to negative log likelihood on PIQA, ARC Easy, SciQ, and OpenWebMath. Crucially, we found that MixMin mixtures for smaller models improved training of larger models, suggesting that MixMin mixtures may be scale-invariant. When mixing bioassay data to train an XGBoost model, we saw improvements to average precision scores of 0.03 -0.15.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b6fcaa8-a7b8-41cf-aae1-65496939e7f6Builds on9
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguagePhilipp Seidl, Andreu Vall, Sepp Hochreiter, Günter KlambauerICML 2023 · 69 citations
Related papers
- Fast Data Mixture Optimization via Gradient DescentHaoru Tan, Sitong Wu, Yanfeng Chen, Jun Xia et al.ICLR 2026
- RegMix: Data Mixture as Regression for Language Model Pre-trainingQian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng et al.ICLR 2025
- Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian FrameworkThomson Yen, Andrew Siah, Haozhe Chen, C. Guetta et al.NeurIPS 2025 · 12 citations
- MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model MergingJiapeng Wang, Changxin Tian, Kunlong Chen, ziqi liu et al.ICML 2026 · 6 citations
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye et al.ICML 2026 · 3 citations
