Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation
Pengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi, Beijun Shen, Xiaodong Gu
Abstract
Software testing is a critical, yet resource-intensive phase of the software development lifecycle. Search-based approaches typically achieve high coverage but produce tests with low readability, whereas large language model (LLM)-based methods generate more human-readable tests but often suffer from low coverage and compilability. While the majority of research efforts have focused on improving test coverage and readability, comparatively less attention has been paid to enhancing the robustness of bug detection. To address this gap, we propose AdverTest, a novel adversarial framework for LLM-powered test case generation that pairs a test case generation agent (T ) with a mutant generation agent (M): M persistently creates mutants "hacking" the blind spots of T 's current test suite, while T iteratively refines its tests to "kill" the challenging mutants, with the interaction guided by both coverage and mutation scores. Experimental results on Defects4J show that our approach improves fault detection rates by 8.56% over the best existing LLM-based methods and by 50.20% over EvoSuite, while remaining competitive on line and branch coverage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edcb6e4b-38c8-4b61-9471-883672f901cfBuilds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao et al.OOPSLA 2024 · 74 citations
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney et al.ICSE 2023 · 45 citations
- Extracting Concise Bug-Fixing Patches from Human-Written Patches in Version Control SystemsYanjie Jiang, Hui Liu, Nan Niu, Lu Zhang et al.ICSE 2021 · 38 citations
Related papers
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li et al.ICLR 2026 · 4 citations
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 8 citations
- HITS: High-coverage LLM-based Unit Test Generation via Method SlicingZejun Wang, Kaibo Liu, Ge Li, Zhi JinASE 2024 · 29 citations
