AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
Luca Traini, Federico Di Menna, Vittorio Cortellessa
摘要
Performance testing aims at uncovering efficiency issues of software systems. In order to be both effective and practical, the design of a performance test must achieve a reasonable trade-off between result quality and testing time. This becomes particularly challenging in Java context, where the software undergoes a warm-up phase of execution, due to just-in-time compilation. During this phase, performance measurements are subject to severe fluctuations, which may adversely affect quality of performance test results. Both practitioners and researchers have proposed approaches to mitigate this issue. Practitioners typically rely on a fixed number of iterated executions that are used to warm-up the software before starting to collect performance measurements (state-of-practice).
Researchers have developed techniques that can dynamically stop warm-up iterations at runtime (state-of-the-art). However, these approaches often provide suboptimal estimates of the warm-up phase, resulting in either insufficient or excessive warm-up iterations, which may degrade result quality or increase testing time. There is still a lack of consensus on how to properly address this problem. Here, we propose and study an AI-based framework to dynamically halt warm-up iterations at runtime. Specifically, our framework leverages recent advances in AI for Time Series Classification (TSC) to predict the end of the warm-up phase during test execution. We conduct experiments by training three different TSC models on half a million of measurement segments obtained from JMH microbenchmark executions. We find that our framework significantly improves the accuracy of the warm-up estimates provided by state-of-practice and state-of-the-art methods. This higher estimation accuracy results in a net improvement in either result quality or testing time for up to +35.3% of the microbenchmarks. Our study highlights that integrating AI to dynamically estimate the end of the warm-up phase can enhance the cost-effectiveness of Java performance testing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Experimental Evaluation Methodology for the Era of No Steady PerformanceJaromír Antoch, Walter Binder, Lubomír Bulej, François Farquet 等OOPSLA 2026
- Understanding and Finding JIT Compiler Performance BugsZijian Yi, Cheng Ding, August Shi, Milos GligoricOOPSLA 2026
它引用的顶会 Paper6
- Omni-Scale CNNs: a simple and effective kernel size configuration for time series classificationWensi Tang, Guodong Long, Lu Liu, Tianyi Zhou 等ICLR 2022 · 被引用 163 次
- Multi-objectivizing software configuration tuningTao Chen, Miqing LiFSE 2021 · 被引用 40 次
- Dynamically reconfiguring software microbenchmarks: reducing execution time without sacrificing result qualityChristoph Laaber, Stefan Würsten, Harald C. Gall, Philipp LeitnerFSE 2020 · 被引用 36 次
- Towards the use of the readily available tests from the release pipeline as performance tests: are we there yet?Zishuo Ding, Jinfu Chen, Weiyi ShangICSE 2020 · 被引用 33 次
- Faster or Slower? Performance Mystery of Python Idioms Unveiled with Empirical EvidenceZejun Zhang, Zhenchang Xing, Xin Xia, Xiwei Xu 等ICSE 2023 · 被引用 16 次
相关 Paper
- Profiling-Guided Bayesian Optimization of JVM ConfigurationsAbdelrahman Baz, Wing Lam, August ShiISSTA 2026
- Divining Profiler Accuracy: An Approach to Approximate Profiler Accuracy through Machine Code-Level SlowdownHumphrey Burchell, Stefan MarrOOPSLA 2025 · 被引用 3 次
- Reducing Test Runtime by Transforming Test FixturesChengpeng Li, Abdelrahman Baz, August ShiASE 2024 · 被引用 1 次
- LLM4JMH: Studying the Use of LLMs for Generating Java Performance MicrobenchmarksZongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang 等ICSE 2026
- JOSer: Just-In-Time Object Serialization for Heavy Java Serialization WorkloadsChaokun Yang, Pengbo Nie, Ziyi Lin, Weipeng Wang 等ASPLOS 2026
