Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Reza Haf
摘要
Large language models (LLMs) are typically fine-tuned on diverse and extensive datasets sourced from various origins to develop a comprehensive range of skills, such as writing, reasoning, chatting, coding, and more. Each skill has unique characteristics, and these datasets are often heterogeneous and imbalanced, making the fine-tuning process highly challenging. Balancing the development of each skill while ensuring the model maintains its overall performance requires sophisticated techniques and careful dataset curation. In this work, we propose a general, model-agnostic, reinforcement learning framework, Mixture-of-Skills (MoS), that learns to optimize data usage automatically during the fine-tuning process. This framework ensures the optimal comprehensive skill development of LLMs by dynamically adjusting the focus on different datasets based on their current learning state. To validate the effectiveness of MoS, we conduct extensive experiments using three diverse LLM backbones on two widely used benchmarks and demonstrate that MoS substantially enhances model performance. Building on the success of MoS, we propose MoSpec, an adaptation for task-specific fine-tuning, which harnesses the utilities of various datasets for a specific purpose. Our work underlines the significance of dataset rebalancing and present MoS as a powerful, general solution for optimizing data usage in the fine-tuning of LLMs for various purposes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language ModelsWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchICLR 2026 · 被引用 2 次
- A Tale of Two Problems: Multi-Task Bilevel Learning Meets Equality Constrained Multi-Objective OptimizationZhiyao Zhang, Myeung Suk Oh, Zhen Qin, Jiaxiang Li 等ICML 2026 · 被引用 1 次
- Boosting Multi-Domain Fine-Tuning of Large Language Models through Evolving Interactions between SamplesXize Liang, Lin Yang, Jie Wang, Yiyang Lu 等ICML 2025
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
相关 Paper
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li 等ACL 2024 · 被引用 39 次
- IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model AlignmentChenlin Ming, Chendi Qu, Qizhi Pei, Zhuoshi Pan 等ICLR 2026 · 被引用 8 次
- Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task AdaptationChen Dun, Mirian del Carmen Hipolito Garcia, Guoqing Zheng, Ahmed Hassan Awadallah 等AAAI 2025 · 被引用 7 次
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen 等NeurIPS 2025 · 被引用 14 次
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language ModelsYuan Li, Zhengzhong Liu, Eric P. XingICML 2025
