A Strategic Coordination Framework of Small LMs Matches Large LMs in Data Synthesis
Xin Gao, Qizhi Pei, Zinan Tang, Yu Li, Honglin Lin, Jiang Wu, Lijun Wu, Conghui He
Abstract
While data synthesis and distillation are promising strategies to enhance small language models, current approaches heavily rely on Large Language Models (LLMs), which suffer from high computational costs, environmental inefficiency, and potential biases inherited from monolithic architectures. In contrast, smaller LMs are more accessible and sustainable, but their individual capabilities often fall short in generating high-quality, diverse, and reliable data. Inspired by collaborative human processes (e.g., peer review), we propose a multiple small LMs involved framework, GRA, that aggregates specialized roles across small LMs to iterative refinement and quality control typically achieved by a single large LM. In this collaborative framework, multiple small LMs assume distinct roles-Generator, Reviewer, and Adjudicator-to simulate a peer-reviewinspired data synthesis pipeline. The Generator proposes initial data samples, the Reviewer critiques their quality and diversity, and the Adjudicator resolves conflicts to finalize the output. By decomposing the synthesis process into specialized sub-tasks, collaborative small LMs can achieve data-level parity with distillation from large LMs. Through experiments across multiple benchmarks, we demonstrate that GRA-produced data matches or exceeds the quality of single large LM outputs, e.g., Qwen-2.5-72B-Instruct. Our results challenge the necessity of monolithic large models for high-quality data synthesis, advocating instead for strategic coordination of smaller agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41588dfc-c10d-4dbb-8cbd-07622ae1cd95Cited by top-tier papers2
- Prompt Optimization with Minimal Unlabeled Input via Meta-ReasoningYuran Sun, Chuan WuICML 2026
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop LearningZinan Tang, Xin Gao, Qizhi Pei, Zhuoshi Pan et al.EMNLP 2025
Builds on7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick et al.ICLR 2024 · 174 citations
- LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMsChansung Park, Juyong Jiang, Fan Wang, Sayak Paul et al.ACL 2025 · 12 citations
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi et al.ASE 2024 · 6 citations
Related papers
- Small But Funny: A Feedback-Driven Approach to Humor DistillationSahithya Ravi, Patrick Huber, Akshat Shrivastava, Vered Shwartz et al.ACL 2024
- Collaborative Enhancement of Large and Small Models for Question Answering via Dual Knowledge TransferShaofei Wang, Yunan Liu, Xiaolan Tang, Wenlong ChenAAAI 2026
- Improving Large Vision and Language Models by Learning from a Panel of PeersJefferson Hernandez, Jing Shi, Simon Jenni, Vicente Ordonez et al.ICCV 2025
- Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural ReasoningChan Young Park, Jillian Fisher, Marius Memmel, Dipika Khullar et al.EMNLP 2025 · 3 citations
- MACoT: Synthesizing Chains of Thought for Small Models via Multi-Agent CollaborationGuokai Tang, Feng ZhaoAAAI 2026
