Validation on machine reading comprehension software without annotated labels: a property-based method
Songqiang Chen, Shuo Jin, Xiaoyuan Xie
摘要
Machine Reading Comprehension (MRC) in Natural Language Processing has seen great progress recently. But almost all the current MRC software is validated with a reference-based method, which requires well-annotated labels for test cases and tests the software by checking the consistency between the labels and the outputs. However, labeling test cases of MRC could be very costly due to their complexity, which makes reference-based validation hard to be extensible and sufficient. Furthermore, solely checking the consistency and measuring the overall score may not be sensible and flexible for assessing the language understanding capability. In this paper, we propose a property-based validation method for MRC software with Metamorphic Testing to supplement the reference-based validation. It does not refer to the labels and hence can make much data available for testing. Besides, it validates MRC software against various linguistic properties to give a specific and in-depth picture on linguistic capabilities of MRC software. Comprehensive experimental results show that our method can successfully reveal violations to the target linguistic properties without the labels. Moreover, it can reveal problems that have been concealed by the traditional validation. Comparison according to the properties provides deeper and more concrete ideas about different language understanding capabilities of the MRC software.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 被引用 37 次
- AEON: a method for automatic evaluation of NLP test casesJen-tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He 等ISSTA 2022 · 被引用 18 次
- COSTELLO: Contrastive Testing for Embedding-Based Large Language Model as a Service EmbeddingsWeipeng Jiang, Juan Zhai, Shiqing Ma, Xiaoyu Zhang 等FSE 2024 · 被引用 1 次
- The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to MistakesMarina Mancoridis, Zoe HitzigICML 2026
相关 Paper
- Testing Your Question Answering Software via Asking RecursivelySongqiang Chen, Shuo Jin, Xiaoyuan XieASE 2021 · 被引用 37 次
- Natural Test Generation for Precise Testing of Question Answering SoftwareQingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang 等ASE 2022 · 被引用 27 次
- Semantics Altering Modifications for Evaluating Comprehension in Machine ReadingViktor Schlegel, Goran Nenadic, Riza Batista-NavarroAAAI 2021 · 被引用 20 次
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao 等ISSTA 2022 · 被引用 24 次
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling AnalysisCongying Xu, Hengcheng Zhu, Songqiang Chen, Jiarong Wu 等FSE 2026 · 被引用 1 次
