CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions
Jingwei Shi, Xinxiang Yin, Jing Huang, Shengyu Tao, Jinman Zhao
Abstract
The evaluation of Large Language Models (LLMs) for code generation relies heavily on the quality and robustness of test cases. However, existing benchmarks often lack coverage for subtle corner cases, allowing incorrect solutions to pass. To bridge this gap, we propose CodeHacker, an automated agent framework dedicated to generating targeted adversarial test cases that expose latent vulnerabilities in program submissions. Mimicking the hack mechanism in competitive programming, CodeHacker employs a multistrategy approach, including stress testing, anti-hash attacks, and logic-specific targeting to break specific code submissions. To ensure the validity and reliability of these attacks, we introduce a Calibration Phase, where the agent iteratively refines its own Validator and Checker via self-generated adversarial probes before evaluating contestant code. Experiments demonstrate that CodeHacker significantly improves the True Negative Rate (TNR) of existing datasets, effectively filtering out incorrect solutions that were previously accepted. Furthermore, generated adversarial cases prove to be superior training data, boosting the performance of RL-trained models on benchmarks like LiveCodeBench. Our code is available at https://github.com/shi0712/ CodeHacker .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c10b715-bef3-4cd8-939b-87f23c1bb838Cited by top-tier papers2
- From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository LevelJia Li, Yuxin Su, Michael R. LyuACL 2026 · 4 citations
- Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and RefinementZexu Sun, Yongcheng Zeng, Erxue Min, Heyang Gao et al.ACL 2026
Builds on5
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible ProgramsKaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang et al.ACL 2025 · 20 citations
- Rethinking Verification for LLM Code Generation: From Generation to TestingZihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu et al.NeurIPS 2025 · 19 citations
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang et al.ACL 2026 · 5 citations
- From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository LevelJia Li, Yuxin Su, Michael R. LyuACL 2026 · 4 citations
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li et al.ICLR 2025
Related papers
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li et al.ICLR 2026 · 4 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr et al.ICML 2025
- Scaling Agentic Verifier for Competitive CodingZeyao Ma, Jing Zhang, Xiaokang Zhang, Jiaxi Yang et al.ICML 2026 · 2 citations
- AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree SearchQingyao Li, Weiwen Liu, Weinan Zhang, Yong Yu et al.ICML 2026
