Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
Zhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen, Yaliang Li, Ying Shen
Abstract
Fine-tuning large language models (LLMs) using diverse datasets is crucial for enhancing their overall performance across various domains. In practical scenarios, existing methods based on modeling the mixture proportions of data composition often struggle with data whose domain labels are missing, imprecise or non-normalized, while methods based on data selection usually encounter difficulties in balancing multi-domain performance. To address these challenges, in this work, we investigate the role of data diversity in enhancing the overall abilities of LLMs by empirically constructing contrastive data pools and theoretically deriving explanations. Building upon the insights gained, we propose a new method that gives the LLM a dual identity: an output model to cognitively probe and select data based on diversity reward, as well as an input model to be tuned with the selected data. Extensive experiments show that the proposed method notably boosts performance across domain-undetermined data and a series of foundational downstream tasks when applied to various advanced LLMs. We release our code and hope this study can shed light on the understanding of data diversity and advance feedback-driven data-model co-design for LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a31a5b93-0c1c-4146-8500-76bc047455eaCited by top-tier papers4
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement FinetuningQianli Shen, Daoyuan Chen, Yilun Huang, Zhenqing Ling et al.ICLR 2026 · 15 citations
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang et al.NeurIPS 2025 · 11 citations
- Exploring Diverse Generation Paths via Inference-time Stiefel Activation SteeringDongxuan Zhu, Ly Tran Ho Khanh, Andy Yat-Ming Cheung, Man-Chung Yue et al.ICLR 2026 · 4 citations
- Grounded in Reality: Learning and Deploying Proactive LLM from Offline LogsFei Wei, Daoyuan Chen, Ce Wang, Yilun Huang et al.ICML 2026 · 2 citations
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu et al.ICLR 2024 · 558 citations
Related papers
- Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationFanqi Wan, Xinting Huang, Tao Yang, Xiaojun Quan et al.EMNLP 2023 · 2 citations
- Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-trainingZheheng Luo, Xin Zhang, Xiao Liu, Haoling Li et al.ACL 2025 · 8 citations
- TANDEM: Bi-Level Data Mixture Optimization with Twin NetworksJiaxing Wang, Deping Xiang, Jin Xu, Mingyang Yi et al.NeurIPS 2025 · 3 citations
- HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language ModelsWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchICLR 2026 · 2 citations
- VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMsKeer Lu, Keshi Zhao, Zhuoran Zhang, Zheng Liang et al.EMNLP 2025 · 3 citations
