Evaluating Language Models as Synthetic Data Generators
Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig
Abstract
Given the increasing use of synthetic data in language model (LM) post-training, an LM's ability to generate high-quality data has become nearly as crucial as its ability to solve problems directly. While prior works have focused on developing effective data generation methods, they lack systematic comparison of different LMs as data generators in a unified setting. To address this gap, we propose AGORABENCH, a benchmark that provides standardized settings and metrics to evaluate LMs' data generation abilities. Through synthesizing 1.26 million training instances using 6 LMs and training 99 student models, we uncover key insights about LMs' data generation capabilities. First, we observe that LMs exhibit distinct strengths. For instance, GPT-4o excels at generating new problems, while Claude-3.5-Sonnet performs better at enhancing existing ones. Furthermore, our analysis reveals that an LM's data generation ability doesn't necessarily correlate with its problem-solving ability. Instead, multiple intrinsic features of data quality-including response quality, perplexity, and instruction difficulty-collectively serve as better indicators. Finally, we demonstrate that strategic choices in output format and costconscious model selection significantly impact data generation effectiveness. Our code, checkpoints, and data are all publicly available at https://github.com/neulab/data-agora . What is the most effective data generation method when using Llama-3? GPT-3 Instruct GPT Chat GPT GPT-4 Llama-3 70B Quality Enhancement Data Generation Methods Data Generator Data Generator Data Generation Methods Instance Generation Response Generation Is there a big difference between using GPT-4o and GPT-4o-mini as data generators? Self Instruct Alpaca Conventional Setting AgoraBench (Ours) Wizard LM Orca Magpie Llama-3.1 8,70,405B GPT-4o Comparable Not Comparable GPT-4o mini
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cccc72d-4ba4-4b1e-bc0f-90ecd83daa33Cited by top-tier papers3
- Optimizing Diversity and Quality through Base-Aligned Model CollaborationYichen Wang, Chenghao Yang, Tenghao Huang, Muhao Chen et al.ICML 2026 · 7 citations
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAChanJoo Jung, Jaehyung KimICLR 2026 · 2 citations
- Your Keywords Know Each Other: Breaking SSE with <1% Leaked DocumentsMingyu Bian, Jiabei Wang, Dandan Xu, Guangyu Huang et al.USENIX Security 2026
Builds on17
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
Related papers
- DataGen: Unified Synthetic Dataset Generation via Large Language ModelsYue Huang, Siyuan Wu, Chujie Gao, Dongping Chen et al.ICLR 2025
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang et al.ICML 2025
- GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang et al.ACL 2026
