AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
Ernie Chang, Yang Li, Patrick Huber, Vish Vogeti, David Kant, Yangyang Shi, Vikas Chandra
Abstract
In language model training, it is desirable to equip models with capabilities from various tasks. However, it is not clear how to directly obtain the right data mixtures for these capabilities as the relationship between data and tasks is difficult to be modeled. In this work, we observe that checkpoint models exhibit emerging capabilities at different points in the training trajectory. Often, the training process saves checkpoints as artifacts that are under-utilized as a source of in-training data signals. We identify these artifact models based on their respective capabilities on the benchmarks and leverage them as data mixers by using their aggregated first-order influence approximation over source data (See Figure 1 ). We demonstrated on eight reasoning benchmarks that the proposed framework shows significant improvements in the pretraining setting, with performance improvements of up to 1.93%. Overall, this shows the potential of checkpoint models to enhance data quality and optimize data mixtures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d1052d2-6f1b-423a-a843-d211cd0e2fceCited by top-tier papers2
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training RecipesChangsheng Zhao, Ernie Chang, Zechun Liu, Chia-Jung Chang et al.ICLR 2026 · 10 citations
- D: Dynamic Directional Graph-Constrained Data Scheduling for LLM TrainingYuanjian Xu, Jianing Hao, Guang Zhang, Zhong LiICML 2026
Builds on8
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- First is Better Than Last for Language Data InfluenceChih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu et al.NeurIPS 2022 · 39 citations
Related papers
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training DataSyeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary et al.ICLR 2026 · 27 citations
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in ReasoningWang Yang, Zirui Liu, Hongye Jin, Qingyu Yin et al.NeurIPS 2025 · 7 citations
- At Which Training Stage Does Code Data Help LLMs Reasoning?Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang et al.ICLR 2024 · 106 citations
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-trainingKailai Yang, Xiao Liu, Lei Ji, Hao Li et al.ACL 2026 · 3 citations
- What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure CodeYuze Zhao, Junpeng Fang, Lu Yu, Zhenya Huang et al.ICML 2026 · 1 citation
