Aioli: A Unified Optimization Framework for Language Model Data Mixing
Mayee F. Chen, Michael Y. Hu, Nicholas Lourie, Kyunghyun Cho, Christopher Ré
Abstract
Language model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to efficiently learn mixture proportions, ranging from fitting regression models over training runs to dynamically updating proportions throughout training. Surprisingly, we find that no existing method consistently outperforms a simple stratified sampling baseline in terms of average test perplexity. To understand this inconsistency, we unify existing methods into a standard framework, showing they are equivalent to solving a common optimization problem: minimize average loss subject to a method-specific mixing law -- an implicit assumption on the relationship between loss and mixture proportions. This framework suggests that measuring the fidelity of a method's mixing law can offer insights into its performance. Empirically, we find that existing methods set their mixing law parameters inaccurately, resulting in the inconsistent mixing performance we observe. Using this insight, we derive a new online method named Aioli, which directly estimates the mixing law parameters throughout training and uses them to dynamically adjust proportions. Aioli outperforms stratified sampling on 6 out of 6 datasets by an average of 0.27 test perplexity points, whereas existing methods fail to consistently beat stratified sampling, doing up to 6.9 points worse. Moreover, in a practical setting where proportions are learned on shorter runs due to computational constraints, Aioli can dynamically adjust these proportions over the full training run, consistently improving performance over existing methods by up to 12.012 test perplexity points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a02bfe3-eae9-40a3-aeb1-492712ed608cCited by top-tier papers12
- Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling LawsZhixuan Pan, Shaowen Wang, Pengfei Liao, Jian LiNeurIPS 2025 · 15 citations
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen et al.NeurIPS 2025 · 14 citations
- Olmix: A Framework for Data Mixing Throughout LM DevelopmentMayee Chen, Tyler Murray, David Heineman, Matt Jordan et al.ICML 2026 · 9 citations
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur et al.SIGMOD 2026 · 5 citations
- TANDEM: Bi-Level Data Mixture Optimization with Twin NetworksJiaxing Wang, Deping Xiang, Jin Xu, Mingyang Yi et al.NeurIPS 2025 · 3 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Adaptive Data Optimization: Dynamic Sample Selection with Scaling LawsYiding Jiang, Allan Zhou, Zhili Feng, Sadhika Malladi et al.ICLR 2025
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan et al.ICLR 2025
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye et al.ICML 2026 · 3 citations
- RegMix: Data Mixture as Regression for Language Model Pre-trainingQian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng et al.ICLR 2025
- MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model MergingJiapeng Wang, Changxin Tian, Kunlong Chen, ziqi liu et al.ICML 2026 · 6 citations
