How Low Can You Go? The Data-Light SE Challenge
Kishan Kumar Ganguly, Tim Menzies
摘要
Much of Software Engineering (SE) research assumes that progress depends on massive datasets and CPUintensive optimizers. Yet has this assumption been rigorously tested?
The counter-evidence presented in this paper suggests otherwise. For over 100 optimization tasks from recent SE papers (including software configuration, performance tuning, product line engineering, project health forecasting, defect prediction, software testing, software process and cost estimation, and cross-domain generalization datasets), even with just a few dozen labels, very simple methods (e.g., diversity sampling, a minimal Bayesian learner, its distance-based non-parametric variant, or random probes) achieve over 90% of the best reported results. Furthermore, these simple methods perform just as well as more complex state-ofthe-the-art optimizers like SMAC, TPE, DEHB etc. While some tasks would require better outcomes and more sampling, these results seen after a few dozen samples would suffice for many engineering needs (particularly when the goal is rapid and cost-efficient guidance rather than slow and exhaustive optimization).
To say that another ways, at least some SE tasks are better served by lightweight approaches that demand fewer labels and far less computation. We hence propose the data-light challenge: when will a handful of labels suffice for SE tasks? To enable a large-scale investigation of this issue, we contribute (1) a mathematical formalization of labeling, (2) lightweight baseline algorithms, and (3) results on public-domain data showing the conditions under which lightweight methods excel or fail.
For the purposes of open science, our scripts and data are online at https://github.com/KKGanguly/NEO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Efficient Compiler Autotuning via Bayesian OptimizationJunjie Chen, Ningxin Xu, Peiqi Chen, Hongyu ZhangICSE 2021 · 被引用 73 次
- PriorBand: Practical Hyperparameter Optimization in the Age of Deep LearningNeeratyoy Mallik, Edward Bergman, Carl Hvarfner, Danny Stoll 等NeurIPS 2023 · 被引用 50 次
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 被引用 42 次
- Predicting Configuration Performance in Multiple Environments with Sequential Meta-LearningJingzhi Gong, Tao ChenFSE 2024 · 被引用 13 次
相关 Paper
- FRUGAL: Unlocking Semi-Supervised Learning for Software AnalyticsHuy Tu, Tim MenziesASE 2021 · 被引用 9 次
- Understanding the automated parameter optimization on transfer learning for cross-project defect prediction: an empirical studyKe Li, Zilin Xiang, Tao Chen, Shuo Wang 等ICSE 2020 · 被引用 54 次
- Learning Configurations for Data-Driven Multi-Objective OptimizationZhiyang Chen, Hailong Yao, Xia YinICML 2025
- Adaptive Partitioning Schemes for Optimistic OptimizationRaja Sunkara, Ardhendu TripathyICML 2025
- CARBS: Compiler Autotuning via Randomized Biased SearchWei Li, Bin Gao, Weng-Fai WongHPDC 2026
