Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models
Rui Li, Qi Liu, Liyang He, Zheng Zhang, Hao Zhang, Shengyu Ye, Junyu Lu, Zhenya Huang
摘要
Code retrieval aims to identify code from extensive codebases that semantically aligns with a given query code snippet. Collecting a broad and high-quality set of query and code pairs is crucial to the success of this task. However, existing data collection methods struggle to effectively balance scalability and annotation quality. In this paper, we first analyze the factors influencing the quality of function annotations generated by Large Language Models (LLMs). We find that the invocation of intra-repository functions and third-party APIs plays a significant role. Building on this insight, we propose a novel annotation method that enhances the annotation context by incorporating the content of functions called within the repository and information on third-party API functionalities. Additionally, we integrate LLMs with a novel sorting method to address the multi-level function call relationships within repositories. Furthermore, by applying our proposed method across a range of repositories, we have developed the Query4Code dataset. The quality of this synthesized dataset is validated through both model training and human evaluation, demonstrating high-quality annotations. Moreover, cost analysis confirms the scalability of our annotation method. 1 * Corresponding Author. 1 Our Code and Dataset is available at https://github. com/smsquirrel/queryAnnotation def export_nb(nb_path): exporter = PythonExporter() output, res = exporter.from_filename(nb_path) if 'outputs' in res: for filename, content in res['outputs'].items(): savefile(filename, content) return output Code Query How to export the content of a Jupyter Notebook file. Docstring Export content from a Jupyter notebook file. Parameters: -nb_path : The file path of the Jupyter notebook to be exported.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VERSE: Verification-based Self-Play for Code InstructionsHao Jiang, Qi Liu, Rui Li, Yuze Zhao 等AAAI 2025 · 被引用 3 次
- CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li 等ACL 2026 · 被引用 3 次
- CLARC: C/C++ Benchmark for Robust Code SearchKaicheng Wang, Liyan Huang, Weike Fang, Weihang WangICLR 2026 · 被引用 3 次
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang 等AAAI 2025 · 被引用 2 次
- ScholarGEC: Enhancing Controllability of Large Language Model for Chinese Academic Grammatical Error CorrectionZixiao Kong, Xianquan Wang, Shuanghong Shen, Keyu Zhu 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
相关 Paper
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and BeyondWenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang 等ICSE 2026 · 被引用 1 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
- SpecAgent: A Speculative Retrieval and Forecasting Agent for Code CompletionGeorge Ma, Anurag Koul, Qi Chen, Yawen Wu 等ACL 2026 · 被引用 3 次
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong 等ASE 2024 · 被引用 7 次
- In Line with Context: Repository-Level Code Generation via Context InliningChao Hu, Wenhao Zeng, Yuling Shi, Beijun Shen 等FSE 2026
