How Low Can You Go? The Data-Light SE Challenge
Kishan Kumar Ganguly, Tim Menzies
Abstract
Much of Software Engineering (SE) research assumes that progress depends on massive datasets and CPUintensive optimizers. Yet has this assumption been rigorously tested?
The counter-evidence presented in this paper suggests otherwise. For over 100 optimization tasks from recent SE papers (including software configuration, performance tuning, product line engineering, project health forecasting, defect prediction, software testing, software process and cost estimation, and cross-domain generalization datasets), even with just a few dozen labels, very simple methods (e.g., diversity sampling, a minimal Bayesian learner, its distance-based non-parametric variant, or random probes) achieve over 90% of the best reported results. Furthermore, these simple methods perform just as well as more complex state-ofthe-the-art optimizers like SMAC, TPE, DEHB etc. While some tasks would require better outcomes and more sampling, these results seen after a few dozen samples would suffice for many engineering needs (particularly when the goal is rapid and cost-efficient guidance rather than slow and exhaustive optimization).
To say that another ways, at least some SE tasks are better served by lightweight approaches that demand fewer labels and far less computation. We hence propose the data-light challenge: when will a handful of labels suffice for SE tasks? To enable a large-scale investigation of this issue, we contribute (1) a mathematical formalization of labeling, (2) lightweight baseline algorithms, and (3) results on public-domain data showing the conditions under which lightweight methods excel or fail.
For the purposes of open science, our scripts and data are online at https://github.com/KKGanguly/NEO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
- Efficient Compiler Autotuning via Bayesian OptimizationJunjie Chen, Ningxin Xu, Peiqi Chen, Hongyu ZhangICSE 2021 · 73 citations
- PriorBand: Practical Hyperparameter Optimization in the Age of Deep LearningNeeratyoy Mallik, Edward Bergman, Carl Hvarfner, Danny Stoll et al.NeurIPS 2023 · 50 citations
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 42 citations
- Predicting Configuration Performance in Multiple Environments with Sequential Meta-LearningJingzhi Gong, Tao ChenFSE 2024 · 13 citations
Related papers
- FRUGAL: Unlocking Semi-Supervised Learning for Software AnalyticsHuy Tu, Tim MenziesASE 2021 · 9 citations
- Understanding the automated parameter optimization on transfer learning for cross-project defect prediction: an empirical studyKe Li, Zilin Xiang, Tao Chen, Shuo Wang et al.ICSE 2020 · 54 citations
- Learning Configurations for Data-Driven Multi-Objective OptimizationZhiyang Chen, Hailong Yao, Xia YinICML 2025
- Adaptive Partitioning Schemes for Optimistic OptimizationRaja Sunkara, Ardhendu TripathyICML 2025
- CARBS: Compiler Autotuning via Randomized Biased SearchWei Li, Bin Gao, Weng-Fai WongHPDC 2026
