Is Historical Data an Appropriate Benchmark for Reviewer Recommendation Systems? : A Case Study of the Gerrit Community
Ian X. Gauthier, Maxime Lamothe, Gunter Mussbacher, Shane McIntosh
Abstract
Reviewer recommendation systems are used to suggest community members to review change requests. Like several other recommendation systems, it is customary to evaluate recommendations using held out historical data. While historybased evaluation makes pragmatic use of available data, historical records may be: (1) overly optimistic, since past assignees may have been suboptimal choices for the task at hand; or (2) overly pessimistic, since "incorrect" recommendations may have been equal (or even better) choices. In this paper, we empirically evaluate the extent to which historical data is an appropriate benchmark for reviewer recommendation systems. We replicate the CHREV and WLRREC approaches and apply them to 9,679 reviews from the GERRIT open source community. We then assess the recommendations with members of the GERRIT reviewing community using quantitative methods (personalized questionnaires about their comfort level with tasks) and qualitative methods (semi-structured interviews). We find that history-based evaluation is far more pessimistic than optimistic in the context of GERRIT review recommendations. Indeed, while 86% of those who had been assigned to a review in the past felt comfortable handling the review, 74% of those labelled as incorrect recommendations also felt that they would have been comfortable reviewing the changes. This indicates that, on the one hand, when reviewer recommendation systems recommend the past assignee, they should indeed be considered correct. Yet, on the other hand, recommendations labelled as incorrect because they do not match the past assignee may have been correct as well. Our results suggest that current reviewer recommendation evaluations do not always model the reality of software development. Future studies may benefit from looking beyond repository data to gain a clearer understanding of the practical value of proposed recommendations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c985ed8-beb2-4a04-b480-76710d2598c6Builds on1
Related papers
- Modeling Review History for Reviewer Recommendation: A Hypergraph ApproachGuoping Rong, Yifan Zhang, Lanxin Yang, Fuli Zhang et al.ICSE 2022 · 20 citations
- CORMS: a GitHub and Gerrit based hybrid code reviewer recommendation approach for modern code reviewPrahar Pandya, Saurabh TiwariFSE 2022 · 21 citations
- Code Recommendation for Open Source Software DevelopersYiqiao Jin, Yunsheng Bai, Yanqiao Zhu, Yizhou Sun et al.WWW 2023 · 28 citations
- Repeated Builds During Code Review: An Empirical Study of the OpenStack CommunityRungroj Maipradit, Dong Wang, Patanamon Thongtanunam, Raula Gaikovina Kula et al.ASE 2023 · 8 citations
- Mitigating the Risk of Defects and Improving Knowledge Distribution with Code Reviewer RecommendersMohammadali Sefidi Esfahani, Peter C. RigbyFSE 2026
