AEON: a method for automatic evaluation of NLP test cases
Jen-tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He, Yuxin Su, Michael R. Lyu
摘要
Due to the labor-intensive nature of manual test oracle construction, various automated testing techniques have been proposed to enhance the reliability of Natural Language Processing (NLP) software. In theory, these techniques mutate an existing test case (e.g., a sentence with its label) and assume the generated one preserves an equivalent or similar semantic meaning and thus, the same label. However, in practice, many of the generated test cases fail to preserve similar semantic meaning and are unnatural (e.g., grammar errors), which leads to a high false alarm rate and unnatural test cases. Our evaluation study finds that 44% of the test cases generated by the state-of-the-art (SOTA) approaches are false alarms. These test cases require extensive manual checking effort, and instead of improving NLP software, they can even degrade NLP software when utilized in model training. To address this problem, we propose AEON for Automatic Evaluation Of NLP test cases. For each generated test case, it outputs scores based on semantic similarity and language naturalness. We employ AEON to evaluate test cases generated by four popular testing techniques on five datasets across three typical NLP tasks. The results show that AEON aligns the best with human judgment. In particular, AEON achieves the best average precision in detecting semantic inconsistent test cases, outperforming the best baseline metric by 10%. In addition, AEON also has the highest average precision of finding unnatural test cases, surpassing the baselines by more than 15%. Moreover, model training with test cases prioritized by AEON leads to models that are more accurate and robust, demonstrating AEON's potential in improving NLP software. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
- Towards Reasonable Budget Allocation in Untargeted Graph Structure Attacks via Gradient DebiasZihan Liu, Yun Luo, Lirong Wu, Zicheng Liu 等NeurIPS 2022 · 被引用 41 次
- MTTM: Metamorphic Testing for Textual Content Moderation SoftwareWenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang 等ICSE 2023 · 被引用 23 次
- Validating Multimedia Content Moderation Software via Semantic FusionWenxuan Wang, Jingyuan Huang, Chang Chen, Jiazhen Gu 等ISSTA 2023 · 被引用 11 次
- LEAP: Efficient and Automated Test Method for NLP SoftwareMingxuan Xiao, Yan Xiao, Hai Dong, Shunhui Ji 等ASE 2023 · 被引用 7 次
它引用的顶会 Paper25
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
相关 Paper
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Natural Test Generation for Precise Testing of Question Answering SoftwareQingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang 等ASE 2022 · 被引用 27 次
- Towards More Realistic Evaluation for Neural Test Oracle GenerationZhongxin Liu, Kui Liu, Xin Xia, Xiaohu YangISSTA 2023 · 被引用 24 次
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling AnalysisCongying Xu, Hengcheng Zhu, Songqiang Chen, Jiarong Wu 等FSE 2026 · 被引用 1 次
