Quickly generating diverse valid test inputs with reinforcement learning
Sameer Reddy, Caroline Lemieux, Rohan Padhye, Koushik Sen
Abstract
Property-based testing is a popular approach for validating the logic of a program. An effective property-based test quickly generates many diverse valid test inputs and runs them through a parameterized test driver. However, when the test driver requires strict validity constraints on the inputs, completely random input generation fails to generate enough valid inputs. Existing approaches to solving this problem rely on whitebox or greybox information collected by instrumenting the input generator and/or test driver. However, collecting such information reduces the speed at which tests can be executed. In this paper, we propose and study a black-box approach for generating valid test inputs. We first formalize the problem of guiding random input generators towards producing a diverse set of valid inputs. This formalization highlights the role of a guide which governs the space of choices within a random input generator. We then propose a solution based on reinforcement learning (RL), using a tabular, on-policy RL approach to guide the generator. We evaluate this approach, RLCheck, against pure random input generation as well as a state-of-the-art greybox evolutionary algorithm, on four real-world benchmarks. We find that in the same time budget, RLCheck generates an order of magnitude more diverse valid inputs than the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney et al.ICSE 2023 · 45 citations
- JITfuzz: Coverage-guided Fuzzing for JVM Just-in-Time CompilersMingyuan Wu, Minghai Lu, Heming Cui, Junjie Chen et al.ICSE 2023 · 36 citations
- BEDIVFUZZ: Integrating Behavioral Diversity into Generator-based FuzzingHoang Lam Nguyen, Lars GrunskeICSE 2022 · 29 citations
- Baffle: Hiding Backdoors in Offline Reinforcement Learning DatasetsChen Gong, Zhou Yang, Yunpeng Bai, Junda He et al.S&P 2024 · 28 citations
- Guiding Greybox Fuzzing with Mutation TestingVasudev Vikram, Isabella Laybourn, Ao Li, Nicole Nair et al.ISSTA 2023 · 22 citations
Related papers
- Tuning Random Generators: Property-Based Testing as Probabilistic ProgrammingRyan Tjoa, Poorva Garg, Harrison Goldstein, Todd D. Millstein et al.OOPSLA 2025 · 2 citations
- Higher income, larger loan? monotonicity testing of machine learning modelsArnab Sharma, Heike WehrheimISSTA 2020 · 12 citations
- Multicore Environment State Representation for Agent-Directed Test GenerationBruno D. Miranda, Luiz M. V. Pereira, Márcio Castro, Luiz C. V. dos SantosDAC 2025 · 1 citation
- Covering All the Bases: Type-Based Verification of Test Input GeneratorsZhe Zhou, Ashish Mishra, Benjamin Delaware, Suresh JagannathanPLDI 2023 · 7 citations
- Compiler Test-Program Generation via Memoized Configuration SearchJunjie Chen, Chenyao Suo, Jiajun Jiang, Peiqi Chen et al.ICSE 2023 · 19 citations
