SWERank: Software Issue Localization with Code Ranking
Revanth Gangi Reddy, Tarun Suresh, JaeHyeok Doo, Ye Liu, Xuan-Phi Nguyen, Yingbo Zhou, Semih Yavuz, Caiming Xiong, Heng Ji, Shafiq Joty
Abstract
Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language issue description (e.g., bug report, feature request), is a critical yet time-consuming aspect of software development. While recent LLM-based agentic approaches demonstrate promise, they often incur significant latency and cost due to complex multi-step reasoning and relying on closed-source LLMs. Alternatively, traditional code ranking models, typically optimized for query-to-code or code-to-code retrieval, struggle with the verbose and failure-descriptive nature of issue localization queries. To bridge this gap, we introduce SWERANK 1 , an efficient and effective retrieve-and-rerank framework for software issue localization. To facilitate training, we construct SWELOC, a large-scale dataset curated from public GitHub repositories, featuring real-world issue descriptions paired with corresponding code modifications. Empirical results on SWE-Bench-Lite and LocBench show that SWERANK achieves state-of-the-art performance, outperforming both prior ranking models and costly agent-based systems using closed-source LLMs like Claude-3.5. Further, we demonstrate SWE-LOC's utility in enhancing various existing retriever and reranker models for issue localization, establishing the dataset as a valuable resource for the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dba57443-bcb0-4302-bd40-85398fd2dd45Cited by top-tier papers3
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric DomainsAustin Xu, Xuan-Phi Nguyen, Yilun Zhou, Chien-Sheng Wu et al.ICLR 2026 · 8 citations
- GraphLocator: Graph-Guided Causal Reasoning for Issue LocalizationWei Liu, Chao Peng, Pengfei Gao, Aofan Liu et al.FSE 2026
- Can Old Tests Do New Tricks for Resolving SWE Issues?Yang Chen, Toufique Ahmed, Reyhaneh Jabbarvand, Martin HirzelFSE 2026
Builds on15
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui et al.EMNLP 2023 · 339 citations
Related papers
- OrcaLoca: An LLM Agent Framework for Software Issue LocalizationZhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang et al.ICML 2025
- Enhancing Issue Localization Agent with Tool-Interactive TrainingZexiong Ma, Chao Peng, Qunhong Zeng, Pengfei Gao et al.ICSE 2026
- LocAgent: Graph-Guided LLM Agents for Code LocalizationZhaoling Chen, Robert Tang, Gangda Deng, Fang Wu et al.ACL 2025
- Towards Explorative IRBL: Combining Semantic Retrieval with LLM-Driven Iterative Code ExplorationMoumita Asad, Rafed Muhammad Yasir, Sam MalekISSTA 2026
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen et al.NeurIPS 2025 · 4 citations
