Who Judges the Judge: An Empirical Study on Online Judge Tests
Kaibo Liu, Yudong Han, Jie M. Zhang, Zhenpeng Chen, Federica Sarro, Mark Harman, Gang Huang, Yun Ma
Abstract
Online Judge platforms play a pivotal role in education, competitive programming, recruitment, career training, and large language model training. They rely on predefined test suites to judge the correctness of submitted solutions. It is therefore important that the solution judgement is reliable and free from potentially misleading false positives (i.e., incorrect solutions that are judged as correct). In this paper, we conduct an empirical study of 939 coding problems with 541,552 solutions, all of which are judged to be correct according to the test suites used by the platform, finding that 43.1% of the problems include false positive solutions (3,700 bugs are revealed in total). We also find that test suites are, nevertheless, of high quality according to widely-studied test effectiveness measurements: 88.2% of false positives have perfect (100%) line coverage, 78.9% have perfect branch coverage, and 32.5% have a perfect mutation score. Our findings indicate that more work is required to weed out false positive solutions and to further improve test suite effectiveness. We have released the detected false positive solutions and the generated test inputs to facilitate future research. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cc49fcd-88b7-47dd-8642-2b0be1cf1319Builds on9
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang et al.FSE 2020 · 121 citations
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 114 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
- Improving Machine Translation Systems via Isotopic ReplacementZeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman et al.ICSE 2022 · 42 citations
- Understanding build issue resolution in practice: symptoms and fix patternsYiling Lou, Zhenpeng Chen, Yanbin Cao, Dan Hao et al.FSE 2020 · 39 citations
Related papers
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Characterizing and Mitigating False-Positive Bug Reports in the Linux KernelJiashuo Tian, Dong Wang, Chen Yang, Haichi Wang et al.FSE 2026
- Automated test generation for REST APIs: no time to rest yetMyeongsoo Kim, Qi Xin, Saurabh Sinha, Alessandro OrsoISSTA 2022 · 67 citations
- Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)Junda Zhao, Shurui Zhou, Eldan CohenISSTA 2026
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
