NaNofuzz: A Usable Tool for Automatic Test Generation
Matthew C. Davis, Sangheon Choi, Sam Estep, Brad A. Myers, Joshua Sunshine
Abstract
In the United States alone, software testing labor is estimated to cost $48 billion USD per year. Despite widespread test execution automation and automation in other areas of software engineering, test suites continue to be created manually by software engineers. We have built a test generation tool, called NaNofuzz, that helps users find bugs in their code by suggesting tests where the output is likely indicative of a bug, e.g., that return NaN (not-a-number) values. NaNofuzz is an interactive tool embedded in a development environment to fit into the programmer's workflow. NaNofuzz tests a function with as little as one button press, analyses the program to determine inputs it should evaluate, executes the program on those inputs, and categorizes outputs to prioritize likely bugs. We conducted a randomized controlled trial with 28 professional software engineers using NaNofuzz as the intervention treatment and the popular manual testing tool, Jest, as the control treatment. Participants using NaNofuzz on average identified bugs more accurately (p < .05, by 30%), were more confident in their tests (p < .03, by 20%), and finished their tasks more quickly (p < .007, by 30%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d574cd5-6986-4675-8ea1-664f77678df4Cited by top-tier papers3
- Rug: Turbo Llm for Rust Unit Test GenerationXiang Cheng, Fan Sang, Yizhuo Zhai, Xiaokuan Zhang et al.ICSE 2025 · 6 citations
- TerzoN: Human-in-the-Loop Software Testing with a Composite OracleMatthew C. Davis, Amy Wei, Brad A. Myers, Joshua SunshineFSE 2025 · 2 citations
- Mock Deep Testing: Toward Separate Development of Data and Models for Deep LearningRuchira Manke, Mohammad Wardat, Foutse Khomh, Hridesh RajanICSE 2025
Builds on3
- UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating FuzzersYuwei Li, Shouling Ji, Yuan Chen, Sizhuang Liang et al.USENIX Security 2021 · 142 citations
- Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3James Bornholt, Rajeev Joshi, Vytautas Astrauskas, Brendan Cully et al.SOSP 2021 · 63 citations
- DeepTC-Enhancer: Improving the Readability of Automatically Generated TestsDevjeet Roy, Ziyi Zhang, Maggie Ma, Venera Arnaoudova et al.ASE 2020 · 32 citations
Related papers
- A Qualitative Analysis of Fuzzer Usability and ChallengesYunze Zhao, Wentao Guo, Harrison Goldstein, Daniel Votipka et al.CCS 2025
- UTopia: Automatic Generation of Fuzz Driver using Unit TestsBokdeuk Jeong, Joonun Jang, Hayoon Yi, Jiin Moon et al.S&P 2023
- A Usability Evaluation of AFL and libFuzzer with CS StudentsStephan Plöger, Mischa Meier, Matthew SmithCHI 2023 · 9 citations
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 8 citations
- ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language ModelDawei Wang, Geng Zhou, Li Chen, Dan Li et al.CCS 2024 · 9 citations
