TerzoN: Human-in-the-Loop Software Testing with a Composite Oracle
Matthew C. Davis, Amy Wei, Brad A. Myers, Joshua Sunshine
摘要
Software testing is difficult, tedious, and may consume 28%-50% of software engineering labor. Automatic test generators aim to ease this burden but have important trade-offs. Fuzzers use an implicit oracle that can detect obviously invalid results, but the oracle problem has no general solution, and an implicit oracle cannot automatically evaluate correctness. Test suite generators like EvoSuite use the program under test as the oracle and therefore cannot evaluate correctness. Property-based testing tools evaluate correctness, but users have difficulty coming up with properties to test and understanding whether their properties are correct. Consequently, practitioners create many test suites manually and often use an example-based oracle to tediously specify correct input and output examples. To help bridge the gaps among various oracle and tool types, we present the Composite Oracle, which organizes various oracle types into a hierarchy and renders a single test result per example execution. To understand the Composite Oracle’s practical properties, we built TerzoN, a test suite generator that includes a particular instantiation of the Composite Oracle. TerzoN displays all the test results in an integrated view composed from the results of three types of oracles and finds some types of test assertion inconsistencies that might otherwise lead to misleading test results. We evaluated TerzoN in a randomized controlled trial with 14 professional software engineers with a popular industry tool, fast-check, as the control. Participants using TerzoN elicited 72% more bugs ( p < 0.01), accurately described more than twice the number of bugs ( p < 0.01) and tested 16% more quickly ( p < 0.05) relative to fast-check.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating FuzzersYuwei Li, Shouling Ji, Yuan Chen, Sizhuang Liang 等USENIX Security 2021 · 被引用 142 次
- A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and ChallengesJenny T. Liang, Chenyang Yang, Brad A. MyersICSE 2024 · 被引用 126 次
- Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3James Bornholt, Rajeev Joshi, Vytautas Astrauskas, Brendan Cully 等SOSP 2021 · 被引用 63 次
- CAT-LM Training Language Models on Aligned Code And TestsNikitha Rao, Kush Jain, Uri Alon, Claire Le Goues 等ASE 2023 · 被引用 37 次
- DeepTC-Enhancer: Improving the Readability of Automatically Generated TestsDevjeet Roy, Ziyi Zhang, Maggie Ma, Venera Arnaoudova 等ASE 2020 · 被引用 32 次
相关 Paper
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- NaNofuzz: A Usable Tool for Automatic Test GenerationMatthew C. Davis, Sangheon Choi, Sam Estep, Brad A. Myers 等FSE 2023 · 被引用 9 次
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 被引用 8 次
- Fuzzing Class SpecificationsFacundo Molina, Marcelo d'Amorim, Nazareno AguirreICSE 2022 · 被引用 25 次
