Example Quality Matters: Multi-Aspects Example Augmentation for Private Library Programming
Yuhao Li, Haifeng Sun, Xuesong Zhang, Shu Yao, Haoyu Zheng, Yvchuan Wang, Huazheng Wang, Zirui Zhuang, Qi Qi, Jianxin Liao, Jingyu Wang
摘要
Recent advances in large language models (LLMs) have significantly improved codegeneration capabilities, particularly through retrieval-augmented generation (RAG) for private libraries. While RAG leverages API documentation to address the scarcity of private code corpora, its performance critically depends on the quality of retrieved examples. Existing approaches often overlook the intrinsic characteristics of these examples, particularly how factors such as complexity, readability, and correctness impact their effectiveness. In this study, we systematically investigate these three critical aspects-complexity, readability, and correctness-and find that optimal examples should exhibit moderate complexity, semantic correctness, and step-by-step execution patterns. Based on these findings, we propose ComboPrompt, a novel example enhancement method that strategically combines existing API examples to improve complexity, refines code structure for readability, and incorporates automated validation ensuring correctness. Extensive evaluations across five private library benchmarks and different LLMs demonstrate that ComboPrompt achieves up to 22% accuracy improvement over baseline approaches. Code is available at GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo 等ACL 2022 · 被引用 208 次
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung 等EMNLP 2023 · 被引用 110 次
相关 Paper
- SRACG: A Code Generation Framework with Selective Retrieval AugmentationMengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang 等AAAI 2026
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and BeyondWenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang 等ICSE 2026 · 被引用 1 次
- Repository-Level Prompt Generation for Large Language Models of CodeDisha Shrivastava, Hugo Larochelle, Daniel TarlowICML 2023 · 被引用 184 次
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 被引用 95 次
- SpecAgent: A Speculative Retrieval and Forecasting Agent for Code CompletionGeorge Ma, Anurag Koul, Qi Chen, Yawen Wu 等ACL 2026 · 被引用 3 次
