Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
Hang Zhou, Yehui Tang, Haochen Qin, Yujie Yang, Renren Jin, Deyi Xiong, Kai Han, Yunhe Wang
Abstract
The efficacy of large language models (LLMs) on downstream tasks usually hinges on instruction tuning, which relies critically on the quality of training data. Unfortunately, collecting high-quality and diverse data is both expensive and time-consuming. To mitigate this issue, we propose a novel Star-Agents framework, which automates the enhancement of data quality across datasets through multi-agent collaboration and assessment. The framework adopts a three-pronged strategy. It initially generates diverse instruction data with multiple LLM agents through a bespoke sampling method. Subsequently, the generated data undergo a rigorous evaluation using a dual-model method that assesses both difficulty and quality. Finaly, the above process evolves in a dynamic refinement phase, where more effective LLMs are prioritized, enhancing the overall data quality. Our empirical studies, including instruction tuning experiments with models such as Pythia and LLaMA, demonstrate the effectiveness of the proposed framework. Optimized datasets have achieved substantial improvements, with an average increase of 12% and notable gains in specific metrics, such as a 40% improvement in Fermi, as evidenced by benchmarks like MT-bench, Vicuna bench, and WizardLM testset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0eb3419-540a-43ce-98be-33969eb5541eCited by top-tier papers7
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
- Synthesizing Post-Training Data for LLMs through Multi-Agent SimulationShuo Tang, Xianghe Pang, Zexi Liu, Bohan Tang et al.ACL 2025 · 22 citations
- Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI CollaborationSirui Han, Yuyao Zhang, Yidan Huang, Xueyan Li et al.CHI 2026 · 3 citations
- Pedagogically-Inspired Data Synthesis for Language Model Knowledge DistillationBowei He, Yankai Chen, Xiaokun Zhang, Linghe Kong et al.ICLR 2026 · 2 citations
- TRACE: Trajectory-based Activation Change Estimation for Task-specific Data SelectionYe He, Shangzhan Li, Yuxin Zhou, Qi ShiAAAI 2026
Builds on17
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
Related papers
- From Selection to Refinement: Iterative Optimization for Instruction DataHang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou et al.ACL 2026
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu et al.EMNLP 2025
- AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task GenerationMengkang Hu, Pu Zhao, Can Xu, Qingfeng Sun et al.KDD 2025 · 6 citations
- Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement LearningGuanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin et al.SIGIR 2026 · 1 citation
- Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction TuningJian Yang, Wei Zhang, Yibo Miao, Shanghaoran Quan et al.ACL 2025 · 4 citations
