API-Guided Dataset Synthesis to Finetune Large Code Models
Zongjie Li, Daoyuan Wu, Shuai Wang, Zhendong Su
摘要
Large code models (LCMs), pre-trained on vast code corpora, have demonstrated remarkable performance across a wide array of code-related tasks. Supervised fine-tuning (SFT) plays a vital role in aligning these models with specific requirements and enhancing their performance in particular domains. However, synthesizing high-quality SFT datasets poses a significant challenge due to the uneven quality of datasets and the scarcity of domain-specific datasets.
Inspired by APIs as high-level abstractions of code that encapsulate rich semantic information in a concise structure, we propose DataScope, an API-guided dataset synthesis framework designed to enhance the SFT process for LCMs in both general and domain-specific scenarios. DataScope comprises two main components: Dsel and Dgen. On one hand, Dsel employs API coverage as a core metric, enabling efficient dataset synthesis in general scenarios by selecting subsets of existing (uneven-quality) datasets with higher API coverage. On the other hand, Dgen recasts domain dataset synthesis as a process of using API-specified high-level functionality and deliberately-constituted code skeletons to synthesize concrete code.
Extensive experiments demonstrate DataScope's effectiveness, with models fine-tuned on its synthesized datasets outperforming those tuned on unoptimized datasets five times larger. Furthermore, a series of analyses on model internals, relevant hyperparameters, and case studies provide additional evidence for the efficacy of our proposed methods. These findings underscore the significance of dataset quality in SFT and advance the field of LCMs by providing an efficient, cost-effective framework for constructing high-quality datasets. This contribution enhances performance across both general and domain-specific scenarios, paving the way for more powerful and tailored LCMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace PartitioningSheng Wang, Pengan Chen, Jingqi Zhou, Qintong Li 等NeurIPS 2025 · 被引用 9 次
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li 等ISSTA 2025 · 被引用 8 次
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMsZongjie Li, Daoyuan Wu, Shuai Wang, Zhendong SuCCS 2025 · 被引用 1 次
- Vul-R2: A Reasoning LLM for Automated Vulnerability RepairXin-Cheng Wen, Zirui Lin, Yijun Yang, Cuiyun Gao 等ASE 2025 · 被引用 1 次
- GraphSynth: Resolving the Diversity-Reliability Trade-off with Probabilistic Factor GraphsZehua Cheng, Wei Dai, Jiahao Sun, Thomas LukasiewiczACL 2026
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li 等NDSS 2026 · 被引用 5 次
- UnitCoder: Scalable Code Synthesis from Pre-training CorporaYichuan Ma, Yunfan Shao, Peiji Li, Demin Song 等EMNLP 2025 · 被引用 2 次
- Domain-Specific Data Synthesis for LLMs through Minimal Sufficient Representation LearningTong Ye, Hang Yu, Tengfei Ma, Xuhong Zhang 等KDD 2026 · 被引用 1 次
- Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case StudyMingwei Liu, Zheng Pei, Yanlin Wang, Zihao Wang 等FSE 2026
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
