LLM4JMH: Studying the Use of LLMs for Generating Java Performance Microbenchmarks
Zongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang, Jinfu Chen, Jiahui Geng, Alexander Pretschner, Jens Grossklags, Manfred Hauswirth, Sonja Schimmler
Abstract
Performance regressions pose a significant threat to software reliability. In recent years, smaller-scale performance testing, such as performance microbenchmarks, has been adopted in practice to detect such regressions in an early stage of development. Developing these microbenchmarks is costly, error-prone, and demands specialized expertise. Motivated by the recent progress of large language models (LLMs) in code-related tasks, we study whether LLMs can step in as performance experts to automate the generation of microbenchmarks for performance regression detection. While existing approaches typically depend on functional unit tests to generate performance microbenchmarks, we instead design an LLM-based approach, LLM4JMH, to assess the capability of LLMs in generating reliable and effective performance microbenchmarks directly from source code. We further explore how program analysis techniques, such as static analysis, can enhance the reliability and effectiveness of LLM-generated performance microbenchmarks. Experimental results show that LLM-generated performance microbenchmarks achieve comparable bug detection effectiveness to expert-written microbenchmarks, while reducing overall expected detection latency by up to 52.60% across RxJava, Eclipse Collections, and Zipkin. In the study, the generated tests can successfully identify six out of 11 real-world performance bugs in Apache Flink. These results demonstrate that LLM-based performance microbenchmark generation can automate early performance regression detection in continuous integration pipelines of the software development process and reduce reliance on expert-crafted tests, advancing performance-aware software engineering.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 33a5f7f7-9100-497a-b8d6-46edf00b80b8Related papers
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu et al.ICSE 2026
- Automated Inline Comment Smell Detection and Repair with Large Language ModelsHatice Kübra Çaglar, Semih Çaglar, Eray TüzünASE 2025 · 1 citation
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
- Can Large Language Models Reason about Program Invariants?Kexin Pei, David Bieber, Kensen Shi, Charles Sutton et al.ICML 2023 · 128 citations
- Evaluating LLM-Based Regression Test GenerationJing Liu, Seongmin Lee, Eleonora Losiouk, Marcel BöhmeFSE 2026 · 1 citation
