FRUGAL: Unlocking Semi-Supervised Learning for Software Analytics
Huy Tu, Tim Menzies
Abstract
Standard software analytics often involves having a large amount of data with labels in order to commission models with acceptable performance. However, prior work has shown that such requirements can be expensive, taking several weeks to label thousands of commits, and not always available when traversing new research problems and domains. Unsupervised Learning is a promising direction to learn hidden patterns within unlabelled data, which has only been extensively studied in defect prediction. Nevertheless, unsupervised learning can be ineffective by itself and has not been explored in other domains (e.g., static analysis and issue close time). Motivated by this literature gap and technical limitations, we present FRUGAL, a tuned semi-supervised method that builds on a simple optimization scheme that does not require sophisticated (e.g., deep learners) and expensive (e.g., 100% manually labelled data) methods. FRUGAL optimizes the unsupervised learner's configurations (via a simple grid search) while validating our design decision of labelling just 2.5% of the data before prediction. As shown by the experiments of this paper FRUGAL outperforms the state-of-the-art adoptable static code warning recognizer and issue closed time predictor, while reducing the cost of labelling by a factor of 40 (from 100% to 2.5%). Hence we assert that FRUGAL can save considerable effort in data labelling especially in validating prior work or researching new problems. Based on this work, we suggest that proponents of complex and expensive methods should always baseline such methods against simpler and cheaper alternatives. For instance, a semi-supervised learner like FRUGAL can serve as a baseline to the state-of-the-art software analytics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 503a3a0c-533b-4352-911e-03f20ec00664Related papers
- Early Life Cycle Software Defect Prediction. Why? How?N. C. Shrikanth, Suvodeep Majumder, Tim MenziesICSE 2021 · 23 citations
- How Low Can You Go? The Data-Light SE ChallengeKishan Kumar Ganguly, Tim MenziesFSE 2026 · 5 citations
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 42 citations
- Loupe: End-to-End Learning of Loop Unrolling Heuristics for Abstract InterpretationMaykel Mattar, Michele Alberti, Valentin Perrelle, Salah SadouASE 2025
- Understanding and Bridging the Gap Between Unsupervised Network Representation Learning and Security AnalyticsJiacen Xu, Xiaokui Shu, Zhou LiS&P 2024 · 14 citations
