Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
Mozhi Zhang, Howe Tissue, Lu Wang, Xipeng Qiu
Abstract
The mixture ratio of data from different source domains significantly affects the performance of language models (LM) pretraining. In this paper, we introduce DOMAIN2VEC, a novel approach that decomposes any dataset into a linear combination of several "Meta-Domains", a new concept designed to capture key underlying features of datasets. DOMAIN2VEC maintains a vocabulary of Meta-Domains and uses a Meta-Domain Classifier to decompose any given dataset into a domain vector that corresponds to a distribution over this vocabulary. These domain vectors enable the identification of optimal data mixture ratio for LM pretraining in a training-free manner under the Distribution Alignment Assumption (DA 2 ), which suggests that when the data distribution of the training set and the validation set is more aligned, a lower validation loss is achieved. Moreover, previous work could use DOMAIN2VEC to model the relationship between domain vectors and LM performance, greatly enhancing the scalability of previous methods without retraining as new datasets are introduced. Extensive experiments demonstrate that DOMAIN2VEC finds data mixture ratios that enhance downstream task performance with minimal computational overhead. Specifically, DO-MAIN2VEC achieves the same validation loss on Pile-CC using only 51.5% of the compute required when training on the original mixture of The Pile Dataset. Under equivalent compute budget, DOMAIN2VEC improves downstream performance by an average of 2.72%. DOMAIN2VEC serves as a strong and efficient baseline for data mixture optimization in LM pretraining, offering insights into improving data efficiency in large-scale models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed887973-30be-4eca-88af-7a0837cb6ba2Cited by top-tier papers4
- Capacity-Aware Mixture Law Enables Efficient LLM Data OptimizationJingwei Li, Xinran Gu, Jingzhao ZhangICML 2026 · 1 citation
- Explaining Data Mixing Scaling Lawsrui dai, SHURAN ZHENGICML 2026 · 1 citation
- Removing Noise, not Finding Gold: Quality Filtering for Large-Scale PretrainingThiziri Nait Saada, Louis Béthune, Michal Klein, David Grangier et al.ICML 2026
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li et al.ICML 2025
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
Related papers
- RegMix: Data Mixture as Regression for Language Model Pre-trainingQian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng et al.ICLR 2025
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye et al.ICML 2026 · 3 citations
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language ModelsSamira Abnar, Harshay Shah, Dan Busbridge, Alaaeldin El-Nouby et al.ICML 2025
- DEPT: Decoupled Embeddings for Pre-training Language ModelsAlex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen et al.ICLR 2025
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao et al.EMNLP 2024 · 2 citations
