CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
Jiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li, Liangwei Chen, Chenyang Lyu, Haonan Li, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray
Abstract
Code retrieval is essential in modern software development, as it boosts code reuse and accelerates debugging. However, current benchmarks primarily emphasize functional relevance while neglecting critical dimensions of software quality. Motivated by this gap, we introduce CoQuIR, the first large-scale, multilingual benchmark specifically designed to evaluate quality-aware code retrieval across four key dimensions: correctness, efficiency, security, and maintainability. CoQuIR provides fine-grained quality annotations for 42,725 queries and 134,907 code snippets in 11 programming languages, and is accompanied by two quality-centric evaluation metrics: Pairwise Preference Accuracy and Margin-based Ranking Score. Using CoQuIR, we benchmark 23 retrieval models, covering both open-source and proprietary systems, and find that even top-performing models frequently fail to distinguish buggy or insecure code from their more robust counterparts. Furthermore, we conduct preliminary investigations into training methods that explicitly encourage retrievers to recognize code quality. Using synthetic datasets, we demonstrate promising improvements in quality-aware metrics across various models, without sacrificing semantic relevance. Downstream code generation experiments further validate the effectiveness of our approach. Overall, our work highlights the importance of integrating quality signals into code retrieval systems, laying the groundwork for more trustworthy and robust software development tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CodeMMR: Bridging Natural Language, Code, and Image for Unified RetrievalJiahui Geng, Qing Li, Fengyu Cai, Fakhri KarrayCVPR 2026
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
Builds on17
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang et al.ICLR 2023 · 140 citations
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
Related papers
- CoIR: A Comprehensive Benchmark for Code Information Retrieval ModelsXiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia et al.ACL 2025
- CLARC: C/C++ Benchmark for Robust Code SearchKaicheng Wang, Liyan Huang, Weike Fang, Weihang WangICLR 2026 · 3 citations
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
- CONSIDER: Commonalities and Specialties Driven Multilingual Code Retrieval FrameworkRui Li, Liyang He, Qi Liu, Yuze Zhao et al.AAAI 2024 · 11 citations
- Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence TasksJiawei Lu, Haoye Wang, Zhongxin Liu, Keyu Liang et al.ASE 2024 · 3 citations
