LLM-Assisted Code Cleaning For Training Accurate Code Generators
Naman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph E. Gonzalez, Koushik Sen, Ion Stoica
摘要
Natural language to code generation is an important application area of LLMS and has received wide attention from the community. The majority of relevant studies have exclusively concentrated on increasing the quantity and functional correctness of training sets while disregarding other stylistic elements of programs. More recently, data quality has garnered a lot of interest and multiple works have showcased its importance for improving performance. In this work, we investigate data quality for code and find that making the code more structured and readable leads to improved code generation performance of the system. We build a novel data-cleaning pipeline that uses these principles to transform existing programs by 1.) renaming variables, 2.) modularizing and decomposing complex code into smaller helper sub-functions, and 3.) inserting natural-language based plans via LLM based transformations. We evaluate our approach on two challenging algorithmic code generation benchmarks and find that fine-tuning CODELLAMA-7B on our transformed modularized programs improves the performance by up to 30% compared to fine-tuning on the original dataset. Additionally, we demonstrate improved performance from using a smaller amount of higher-quality data, finding that a model fine-tuned on the entire original dataset is outperformed by a model trained on 15% of our cleaned dataset. Even in comparison to closed-source models, our models outperform the much larger ALPHACODE models (Li et al., 2022) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model ServingYuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray 等SIGCOMM 2024 · 被引用 111 次
- R2E: Turning any Github Repository into a Programming Agent EnvironmentNaman Jain, Manish Shetty, Tianjun Zhang, King Han 等ICML 2024 · 被引用 50 次
- The Limits of Inference Scaling Through ResamplingBenedikt Stroebl, Sayash Kapoor, Arvind NarayananICLR 2026 · 被引用 38 次
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura 等ICLR 2026 · 被引用 21 次
- When to Stop? Towards Efficient Code Generation in LLMs with Excess Token PreventionLianghong Guo, Yanlin Wang, Ensheng Shi, Wanjun Zhong 等ISSTA 2024 · 被引用 16 次
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
相关 Paper
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM BudgetDangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo 等ICSE 2026
- To Code or Not To Code? Exploring Impact of Code in Pre-trainingViraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot 等ICLR 2025 · 被引用 3 次
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataZifan Song, Yudong Wang, Wenwei Zhang, Kuikun Liu 等NeurIPS 2024 · 被引用 8 次
- Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code GenerationSanjeepan Sivapiran, Gias UddinFSE 2026
