Who Judges the Judge: An Empirical Study on Online Judge Tests
Kaibo Liu, Yudong Han, Jie M. Zhang, Zhenpeng Chen, Federica Sarro, Mark Harman, Gang Huang, Yun Ma
摘要
Online Judge platforms play a pivotal role in education, competitive programming, recruitment, career training, and large language model training. They rely on predefined test suites to judge the correctness of submitted solutions. It is therefore important that the solution judgement is reliable and free from potentially misleading false positives (i.e., incorrect solutions that are judged as correct). In this paper, we conduct an empirical study of 939 coding problems with 541,552 solutions, all of which are judged to be correct according to the test suites used by the platform, finding that 43.1% of the problems include false positive solutions (3,700 bugs are revealed in total). We also find that test suites are, nevertheless, of high quality according to widely-studied test effectiveness measurements: 88.2% of false positives have perfect (100%) line coverage, 78.9% have perfect branch coverage, and 32.5% have a perfect mutation score. Our findings indicate that more work is required to weed out false positive solutions and to further improve test suite effectiveness. We have released the detected false positive solutions and the generated test inputs to facilitate future research. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang 等FSE 2020 · 被引用 121 次
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 被引用 114 次
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis 等ICSE 2020 · 被引用 111 次
- Improving Machine Translation Systems via Isotopic ReplacementZeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman 等ICSE 2022 · 被引用 42 次
- Understanding build issue resolution in practice: symptoms and fix patternsYiling Lou, Zhenpeng Chen, Yanbin Cao, Dan Hao 等FSE 2020 · 被引用 39 次
相关 Paper
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du 等ICSE 2026
- Characterizing and Mitigating False-Positive Bug Reports in the Linux KernelJiashuo Tian, Dong Wang, Chen Yang, Haichi Wang 等FSE 2026
- Automated test generation for REST APIs: no time to rest yetMyeongsoo Kim, Qi Xin, Saurabh Sinha, Alessandro OrsoISSTA 2022 · 被引用 67 次
- Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)Junda Zhao, Shurui Zhou, Eldan CohenISSTA 2026
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 被引用 5 次
