Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
Huihui Huang, Ratnadira Widyasari, Ting Zhang, Ivana Clairine Irsan, Jieke Shi, Han Wei Ang, Frank Liauw, Eng Lieh Ouh, Lwin Khin Shar, Hong Jin Kang, David Lo
摘要
Issue-commit linking, which connects issues with commits that fix them, is crucial for software maintenance. Existing approaches have shown promise in automatically recovering these links. Evaluations of these techniques assess their ability to identify genuine links from plausible but false links. However, these evaluations overlook the fact that, in reality, when a repository has more commits, the presence of more plausible yet unrelated commits may interfere with the tool in differentiating the correct fix commits. To address this, we propose the Realistic Distribution Setting (RDS) and use it to construct a more realistic evaluation dataset that includes 20 opensource projects. By evaluating tools on this dataset, we observe that the performance of the state-of-the-art deep learning-based approach drops by more than half, while the traditional Information Retrieval method, VSM, outperforms it.
Inspired by these observations, we propose EasyLink, which utilizes a vector database as a modern Information Retrieval technique. To address the long-standing problem of the semantic gap between issues and commits, EasyLink leverages a large language model to rerank the commits retrieved from the database. Under our evaluation, EasyLink achieves an average Precision@1 of 75.03%, improving over the state-of-the-art by over four times. Additionally, this paper provides practical guidelines for advancing research in issue-commit link recovery.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An 等ACL 2026 · 被引用 5 次
- LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link RecoveryArshia Akhavan, Alireza Hoseinpour, Abbas Heydarnoori, Hamid Bagheri 等FSE 2026
它引用的顶会 Paper11
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang 等ICSE 2021 · 被引用 124 次
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu 等WWW 2025 · 被引用 50 次
相关 Paper
- EALink: An Efficient and Accurate Pre-Trained Framework for Issue-Commit Link RecoveryChenyuan Zhang, Yanlin Wang, Zhao Wei, Yong Xu 等ASE 2023 · 被引用 10 次
- Not Every Patch is an Island: LLM-Enhanced Identification of Multiple Vulnerability PatchesYi Song, Dongchen Xie, Lin Xu, He Zhang 等ASE 2025
- Delving into Commit-Issue Correlation to Enhance Commit Message Generation ModelsLiran Wang, Xunzhu Tang, Yichen He, Changyu Ren 等ASE 2023 · 被引用 11 次
- PatchFinder: A Two-Phase Approach to Security Patch Tracing for Disclosed Vulnerabilities in Open-Source SoftwareKaixuan Li, Jian Zhang, Sen Chen, Han Liu 等ISSTA 2024 · 被引用 8 次
- Semi-supervised pre-processing for learning-based traceability framework on real-world software projectsLiming Dong, He Zhang, Wei Liu, Zhiluo Weng 等FSE 2022 · 被引用 15 次
