Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
Wanyun Xie, Francesco Tonin, Volkan Cevher
Abstract
Training data mixtures greatly impact the generalization performance of large language models. Existing domain reweighting methods often rely on costly weight computations and require retraining when new data is introduced. To this end, we introduce a flexible and efficient data mixing framework, CHAMELEON, that employs leverage scores to quantify domain importance within a learned embedding space. We first construct a domain affinity matrix over domain embeddings. The induced leverage scores determine a mixture that upweights domains sharing common representations in embedding space. This formulation allows direct transfer to new data by computing the new domain embeddings. In experiments, we demonstrate improvements over three key scenarios: (i) our computed weights improve performance on pretraining domains with a fraction of the compute of existing methods; (ii) CHAMELEON can adapt to data changes without proxy retraining, boosting few-shot reasoning accuracies when transferred to new data; (iii) our method enables efficient domain reweighting in finetuning, consistently improving test perplexity on all finetuning domains over uniform mixture. Our code is available at https:// github.com/LIONS-EPFL/Chameleon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 207c6730-ba28-44c4-94c3-0cd462411679Cited by top-tier papers9
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen et al.NeurIPS 2025 · 14 citations
- Olmix: A Framework for Data Mixing Throughout LM DevelopmentMayee Chen, Tyler Murray, David Heineman, Matt Jordan et al.ICML 2026 · 9 citations
- DUET: Optimizing LLM Training Data Mixtures via Noisy Feedback from Unseen, Downstream Evaluation TasksZhiliang Chen, Gregory Kang Ruey Lau, Chuan Sheng Foo, Bryan Kian Hsiang LowICLR 2026 · 8 citations
- Procedural Pretraining: Warming Up Language Models with Abstract DataLiangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran et al.ICML 2026 · 6 citations
- TANDEM: Bi-Level Data Mixture Optimization with Twin NetworksJiaxing Wang, Deping Xiang, Jin Xu, Mingyang Yi et al.NeurIPS 2025 · 3 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
Related papers
- DOGE: Domain Reweighting with Generalization EstimationSimin Fan, Matteo Pagliardini, Martin JaggiICML 2024 · 79 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-trainingKailai Yang, Xiao Liu, Lei Ji, Hao Li et al.ACL 2026 · 3 citations
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language ModelsYuan Li, Zhengzhong Liu, Eric P. XingICML 2025
- Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsYihao Xue, Ali Payani, Yu Yang, Baharan MirzasoleimanICML 2024 · 4 citations
