Natural Test Generation for Precise Testing of Question Answering Software
Qingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang, Shuang Liu, Menghan Tian
Abstract
Question answering (QA) software uses information retrieval and natural language processing techniques to automatically answer questions posed by humans in a natural language. Like other AI-based software, QA software may contain bugs. To automatically test QA software without human labeling, previous work extracts facts from question answer pairs and generates new questions to detect QA software bugs. Nevertheless, the generated questions could be ambiguous, confusing, or with chaotic syntax, which are unanswerable for QA software. As a result, a relatively large proportion of the reported bugs are false positives. In this work, we proposed QAQA, a sentence-level mutation based metamorphic testing technique for QA software. To eliminate false positives and achieve precise automatic testing, QAQA leverages five Metamorphic Relations (MRs) as well as semantics-guided search and enhanced test oracles. Our evaluation on three QA datasets demonstrates that QAQA outperforms the state-of-the-art in both quantity (8,133 vs. 6,601 bugs) and quality (97.67% vs. 49% true positive rate) of the reported bugs. Moreover, the test inputs generated by QAQA successfully reduce MR violation rate from 44.29% to 20.51% when being adopted in fine-tuning the QA software under test.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9d2a834d-945f-41a5-8822-01c68852471bCited by top-tier papers10
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu et al.FSE 2023 · 50 citations
- Regression Fuzzing for Deep Learning SystemsHanmo You, Zan Wang, Junjie Chen, Shuang Liu et al.ICSE 2023 · 28 citations
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- Validating Multimedia Content Moderation Software via Semantic FusionWenxuan Wang, Jingyuan Huang, Chang Chen, Jiazhen Gu et al.ISSTA 2023 · 11 citations
- ChatGPT Incorrectness Detection in Software ReviewsMinaoar Hossain Tanzil, Junaed Younus Khan, Gias UddinICSE 2024 · 10 citations
Related papers
- Testing Your Question Answering Software via Asking RecursivelySongqiang Chen, Shuo Jin, Xiaoyuan XieASE 2021 · 37 citations
- Knowledge Graph Driven Inference Testing for Question Answering SoftwareJun Wang, Yanhui Li, Zhifei Chen, Lin Chen et al.ICSE 2024 · 2 citations
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling AnalysisCongying Xu, Hengcheng Zhu, Songqiang Chen, Jiarong Wu et al.FSE 2026 · 1 citation
- Validation on machine reading comprehension software without annotated labels: a property-based methodSongqiang Chen, Shuo Jin, Xiaoyuan XieFSE 2021 · 27 citations
- MorphQ: Metamorphic Testing of the Qiskit Quantum Computing PlatformMatteo Paltenghi, Michael PradelICSE 2023 · 44 citations
