23 shades of self-admitted technical debt: an empirical study on machine learning software
David O'Brien, Sumon Biswas, Sayem Imtiaz, Rabe Abdalkareem, Emad Shihab, Hridesh Rajan
Abstract
In software development, the term “technical debt” (TD) is used to characterize short-term solutions and workarounds implemented in source code which may incur a long-term cost. Technical debt has a variety of forms and can thus affect multiple qualities of software including but not limited to its legibility, performance, and structure. In this paper, we have conducted a comprehensive study on the technical debts in machine learning (ML) based software. TD can appear differently in ML software by infecting the data that ML models are trained on, thus affecting the functional behavior of ML systems. The growing inclusion of ML components in modern software systems have introduced a new set of TDs. Does ML software have similar TDs to traditional software? If not, what are the new types of ML specific TDs? Which ML pipeline stages do these debts appear? Do these debts differ in ML tools and applications and when they get removed? Currently, we do not know the state of the ML TDs in the wild. To address these questions, we mined 68,820 self-admitted technical debts (SATD) from all the revisions of a curated dataset consisting of 2,641 popular ML repositories from GitHub, along with their introduction and removal. By applying an open-coding scheme and following upon prior works, we provide a comprehensive taxonomy of ML SATDs. Our study analyzes ML SATD type organizations, their frequencies within stages of ML software, the differences between ML SATDs in applications and tools, and quantifies the removal of ML SATDs. The findings discovered suggest implications for ML developers and researchers to create maintainable ML systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 34b41e2f-61c4-47f6-888e-91d3348e2f06Cited by top-tier papers2
- Are Prompt Engineering and TODO Comments Friends or Foes? An Evaluation on GitHub CopilotDavid O'Brien, Sumon Biswas, Sayem Mohammad Imtiaz, Rabe Abdalkareem et al.ICSE 2024 · 7 citations
- The Product Beyond the Model - An Empirical Study of Repositories of Open-Source ML ProductsNadia Nahar, Haoran Zhang, Grace A. Lewis, Shurui Zhou et al.ICSE 2025 · 2 citations
Related papers
- An Empirical Study of Refactorings and Technical Debt in Machine Learning SystemsYiming Tang, Raffi Khatchadourian, Mehdi Bagherzadeh, Rhia Singh et al.ICSE 2021 · 60 citations
- Towards Automatically Addressing Self-Admitted Technical Debt: How Far Are We?Antonio Mastropaolo, Massimiliano Di Penta, Gabriele BavotaASE 2023 · 13 citations
- Detecting and Explaining Self-Admitted Technical Debts with Attention-based Neural NetworksXin Wang, Jin Liu, Li Li, Xiao Chen et al.ASE 2020 · 25 citations
- Evaluating Unit Testing Practices in R PackagesMelina C. VidoniICSE 2021 · 14 citations
- A Large-Scale Study of Model Integration in ML-Enabled Software SystemsYorick Sens, Henriette Knopp, Sven Peldszus, Thorsten BergerICSE 2025 · 3 citations
