CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Hao Zhang, Xinyi Dai, Yasheng Wang, Ruiming Tang
Abstract
Despite the substantial success of Information Retrieval (IR) in various NLP tasks, most IR systems predominantly handle queries and corpora in natural language, neglecting the domain of code retrieval. Code retrieval is critically important yet remains under-explored, with existing methods and benchmarks inadequately representing the diversity of code in various domains and tasks. Moreover, many models have begun to overfit existing leaderboards, limiting their generalizability and real-world applicability. Addressing this gap, we introduce COIR (Code Information Retrieval Benchmark), a robust and comprehensive benchmark specifically designed to evaluate code retrieval capabilities. COIR consists of ten meticulously curated code datasets, all of which have undergone thorough manual inspection and processing. These datasets cover eight distinct retrieval tasks across seven diverse domains, ensuring a broad and rigorous assessment of code retrieval performance. We first discuss the construction of COIR and its diverse dataset composition. Further, we evaluate ten widely used retrieval models using COIR, uncovering significant difficulties in performing code retrieval tasks even with state-of-the-art systems. To ensure seamless integration, COIR is released as a user-friendly Python framework, aligned with the data schema of MTEB and BEIR for consistent cross-benchmark evaluation. Through COIR, we aim to invigorate research in the code retrieval domain, providing a versatile benchmarking tool that encourages further development and exploration of code retrieval systems 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- mmBERT: A Modern Multilingual Encoder with Annealed Language LearningMarc Marone, Orion Weller, William Fleshman, Eugene Yang et al.ICML 2026 · 49 citations
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan et al.ICML 2026 · 15 citations
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid SearchMengzhao Wang, Boyu Tan, Yunjun Gao, Hai Jin et al.VLDB 2026 · 8 citations
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen et al.ICLR 2026 · 3 citations
Builds on5
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung et al.EMNLP 2023 · 110 citations
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang et al.NeurIPS 2025 · 85 citations
- CoSQA: 20, 000+ Web Queries for Code Search and Question AnsweringJunjie Huang, Duyu Tang, Linjun Shou, Ming Gong et al.ACL 2021
- UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationDaya Guo, Shuai Lu, Nan Duan, Yanlin Wang et al.ACL 2022
Related papers
- CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li et al.ACL 2026 · 3 citations
- CodeMMR: Bridging Natural Language, Code, and Image for Unified RetrievalJiahui Geng, Qing Li, Fengyu Cai, Fakhri KarrayCVPR 2026
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi et al.ICLR 2025
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang et al.ACL 2025
- CLARC: C/C++ Benchmark for Robust Code SearchKaicheng Wang, Liyan Huang, Weike Fang, Weihang WangICLR 2026 · 3 citations
