Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation
Pengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi, Beijun Shen, Xiaodong Gu
摘要
Software testing is a critical, yet resource-intensive phase of the software development lifecycle. Search-based approaches typically achieve high coverage but produce tests with low readability, whereas large language model (LLM)-based methods generate more human-readable tests but often suffer from low coverage and compilability. While the majority of research efforts have focused on improving test coverage and readability, comparatively less attention has been paid to enhancing the robustness of bug detection. To address this gap, we propose AdverTest, a novel adversarial framework for LLM-powered test case generation that pairs a test case generation agent (T ) with a mutant generation agent (M): M persistently creates mutants "hacking" the blind spots of T 's current test suite, while T iteratively refines its tests to "kill" the challenging mutants, with the interaction guided by both coverage and mutation scores. Experimental results on Defects4J show that our approach improves fault detection rates by 8.56% over the best existing LLM-based methods and by 50.20% over EvoSuite, while remaining competitive on line and branch coverage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao 等OOPSLA 2024 · 被引用 74 次
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney 等ICSE 2023 · 被引用 45 次
- Extracting Concise Bug-Fixing Patches from Human-Written Patches in Version Control SystemsYanjie Jiang, Hui Liu, Nan Niu, Lu Zhang 等ICSE 2021 · 被引用 38 次
相关 Paper
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li 等ICLR 2026 · 被引用 4 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 被引用 8 次
- HITS: High-coverage LLM-based Unit Test Generation via Method SlicingZejun Wang, Kaibo Liu, Ge Li, Zhi JinASE 2024 · 被引用 29 次
