Understanding and Improving Flaky Test Classification
Shanto Rahman, Saikat Dutta, August Shi
Abstract
Regression testing is an essential part of software development, but it suffers from the presence of flaky tests -tests that pass and fail non-deterministically when run on the same code. These unpredictable failures waste developers' time and often hide real bugs. Prior work showed that fine-tuned large language models (LLMs) can classify flaky tests into different categories with very high accuracy. However, we find that prior approaches over-estimated the accuracy of the models due to incorrect experimental design and unrealistic datasets -making the flaky test classification problem seem simpler than it is.
In this paper, we first show how prior flaky test classifiers over-estimate the prediction accuracy due to 1) flawed experiment design and 2) mis-representation of the real distribution of flaky (and non-flaky) tests in their datasets. After we fix the experimental design and construct a more realistic dataset (which we name FlakeBench), the prior state-of-the-art model shows a steep drop in F1-score, from 81.82% down to 56.62%. Motivated by these observations, we develop a new training strategy to fine-tune a flaky test classifier, FlakyLens, that improves the classification F1-score to 65.79% (9.17pp higher than the state-of-the-art). We also compare FlakyLens against recent pre-trained LLMs, such as CodeLlama and DeepSeekCoder, on the same classification task. Our results show that FlakyLens consistently outperforms these models, highlighting that general-purpose LLMs still fall short on this specialized task.
Using our improved flaky test classifier, we identify the important tokens in the test code that influence the models in making correct or incorrect predictions. By leveraging attribution scores computed per code token in each test, we investigate the tokens that have higher impact on the flaky test classifier's decision-making per flaky test category. To assess the influence of these important tokens, we introduce adversarial perturbation using these important tokens into the tests and observe whether the model's predictions change. Our findings show that, when introducing perturbations using the most important tokens, the classification accuracy can change by as much as -18.37pp. These results highlight that these models still struggle to generalize beyond their training data and rely on identifying category-specific tokens (instead of understanding their semantic context), calling for further research into more robust training methodologies.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62fca171-545e-402c-9bcb-d8b576ab3053Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 150 citations
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang et al.ISSTA 2020 · 71 citations
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 63 citations
- Discretized Integrated Gradients for Explaining Language ModelsSoumya Sanyal, Xiang RenEMNLP 2021 · 34 citations
- FLEX: fixing flaky tests in machine learning projects by updating assertion boundsSaikat Dutta, August Shi, Sasa MisailovicFSE 2021 · 33 citations
Related papers
- Ranking Relevant Tests for Order-Dependent Flaky TestsShanto Rahman, Bala Naren Chanumolu, Suzzana Rafi, August Shi et al.ICSE 2025 · 1 citation
- LLM4JMH: Studying the Use of LLMs for Generating Java Performance MicrobenchmarksZongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang et al.ICSE 2026
- Evolution-aware detection of order-dependent flaky testsChengpeng Li, August ShiISSTA 2022 · 12 citations
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu et al.ISSTA 2025 · 7 citations
- A large-scale longitudinal study of flaky testsWing Lam, Stefan Winter, Anjiang Wei, Tao Xie et al.OOPSLA 2020 · 63 citations
