DataRecipe - How to Cook the Data for CodeLLM?
Kisub Kim, Jounghoon Kim, Byeongjo Park, Dongsun Kim, Chun Yong Chong, Yuan Wang, Tiezhu Sun, Daniel Tang, Jacques Klein, Tegawendé F. Bissyandé
Abstract
Despite the proliferation of language models, a lack of transparency persists regarding the training datasets used. Security concerns are often cited, but identifying high-quality training data is crucial for optimal model performance. Yet, while significant efforts have been made to improve model performance, dataset quality remains an under-explored area. Our study addresses this gap by comprehensively investigating data-quality properties and processing strategies used to train code generation models. We focus on identifying dataset features that impact model performance and leverage these insights to optimize datasets and enhance model efficacy. Our approach involves a multifaceted analysis encompassing metadata, statistics, data quality issues, semantic correlations between intent and code, and design choices. By manipulating these features, we explore their influence on model performance. Our findings reveal that dataset design choices significantly impact the performance of code generation models. Additionally, semantic correlations between intent and code can also affect performance, although to varying degrees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsIrene Solaiman, Christy DennisonNeurIPS 2021 · 276 citations
Related papers
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 1 citation
- LLM-Assisted Code Cleaning For Training Accurate Code GeneratorsNaman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph E. Gonzalez et al.ICLR 2024 · 49 citations
- Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled PreferencesMohammad Saqib Hasan, Saikat Chakraborty, Santu Karmaker, Niranjan BalasubramanianACL 2025
- On Code-Induced Reasoning in LLMsAbdul Waheed, Zhen Wu, Carolyn Rose, Daphne IppolitoICLR 2026 · 6 citations
