CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Hao Zhang, Xinyi Dai, Yasheng Wang, Ruiming Tang
摘要
Despite the substantial success of Information Retrieval (IR) in various NLP tasks, most IR systems predominantly handle queries and corpora in natural language, neglecting the domain of code retrieval. Code retrieval is critically important yet remains under-explored, with existing methods and benchmarks inadequately representing the diversity of code in various domains and tasks. Moreover, many models have begun to overfit existing leaderboards, limiting their generalizability and real-world applicability. Addressing this gap, we introduce COIR (Code Information Retrieval Benchmark), a robust and comprehensive benchmark specifically designed to evaluate code retrieval capabilities. COIR consists of ten meticulously curated code datasets, all of which have undergone thorough manual inspection and processing. These datasets cover eight distinct retrieval tasks across seven diverse domains, ensuring a broad and rigorous assessment of code retrieval performance. We first discuss the construction of COIR and its diverse dataset composition. Further, we evaluate ten widely used retrieval models using COIR, uncovering significant difficulties in performing code retrieval tasks even with state-of-the-art systems. To ensure seamless integration, COIR is released as a user-friendly Python framework, aligned with the data schema of MTEB and BEIR for consistent cross-benchmark evaluation. Through COIR, we aim to invigorate research in the code retrieval domain, providing a versatile benchmarking tool that encourages further development and exploration of code retrieval systems 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- mmBERT: A Modern Multilingual Encoder with Annealed Language LearningMarc Marone, Orion Weller, William Fleshman, Eugene Yang 等ICML 2026 · 被引用 49 次
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan 等ICML 2026 · 被引用 15 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid SearchMengzhao Wang, Boyu Tan, Yunjun Gao, Hai Jin 等VLDB 2026 · 被引用 8 次
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper5
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung 等EMNLP 2023 · 被引用 110 次
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang 等NeurIPS 2025 · 被引用 85 次
- CoSQA: 20, 000+ Web Queries for Code Search and Question AnsweringJunjie Huang, Duyu Tang, Linjun Shou, Ming Gong 等ACL 2021
- UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationDaya Guo, Shuai Lu, Nan Duan, Yanlin Wang 等ACL 2022
相关 Paper
- CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li 等ACL 2026 · 被引用 3 次
- CodeMMR: Bridging Natural Language, Code, and Image for Unified RetrievalJiahui Geng, Qing Li, Fengyu Cai, Fakhri KarrayCVPR 2026
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi 等ICLR 2025
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang 等ACL 2025
- CLARC: C/C++ Benchmark for Robust Code SearchKaicheng Wang, Liyan Huang, Weike Fang, Weihang WangICLR 2026 · 被引用 3 次
