ATGen: Adversarial Reinforcement Learning for Test Case Generation
Qingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang
Abstract
Large Language Models (LLMs) excel at code generation, yet their outputs often contain subtle bugs, for which effective test cases are a critical bottleneck. Existing test generation methods, whether based on prompting or supervised fine-tuning, rely on static datasets. This imposes a “fixed-difficulty ceiling”, fundamentally limiting their ability to uncover novel or more complex bugs beyond their training scope. To overcome this, we introduce ATGEN, a framework that trains a test case generator via adversarial reinforcement learning. ATGEN pits a test generator against an adversarial code generator that continuously crafts harder bugs to evade the current policy. This dynamic loop creates a curriculum of increasing difficulty that continuously challenges the current policy. The test generator is optimized via Reinforcement Learning (RL) to jointly maximize “Output Accuracy” and “Attack Success”, enabling it to learn a progressively stronger policy that breaks the fixed-difficulty ceiling of static training. Extensive experiments demonstrate that ATGEN significantly outperforms state-of-the-art baselines. We further validate its practical utility, showing it serves as both a more effective filter for Best-of-N inference and a higher-quality reward source for training code generation models. Our work establishes a new, dynamic paradigm for improving the reliability of LLM-generated code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b667a3f-0518-4482-9a2c-bab24a8a79d4Builds on6
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang et al.ASE 2024 · 42 citations
Related papers
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 14 citations
- Co-Evolving LLM Coder and Unit Tester via Reinforcement LearningYinjie Wang, Ling Yang, Ye Tian, Ke Shen et al.NeurIPS 2025 · 56 citations
- CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsJingwei Shi, Xinxiang Yin, Jing Huang, Shengyu Tao et al.ACL 2026 · 6 citations
- HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithimic CodingZhongmou He, Yee Man Choi, Kexun Zhang, Ivan Bercovich et al.ICLR 2026
- Themis: Automated Constraint-Aware Test Synthesis Framework for Code Reinforcement LearningShengyu Ye, Qi Liu, Hao Jiang, Zheng Zhang et al.AAAI 2026
