ATOM: Automated Black-Box Testing of Multi-Label Image Classification Systems
Shengyou Hu, Huayao Wu, Peng Wang, Jing Chang, Yongjun Tu, Xiu Jiang, Xintao Niu, Changhai Nie
摘要
Multi-label Image Classification Systems (MICSs) developed based on Deep Neural Networks (DNNs) are extensively used in people's daily life. Currently, although there are a variety of approaches to test DNN-based systems, they typically rely on the internals of DNNs to design test cases, and do not take the core specification of MICS (i.e., correctly recognizing multiple objects in a given image) into account. In this paper, we propose ATOM, an automated and systematic black-box testing framework for testing MICS. Specifically, ATOM exploits the label combination as the testing adequacy criteria, hoping to systematically examine the impact of correlations between a fixed number of labels on the classification ability of MICS. Then, ATOM leverages image search engine and natural language processing to find test images that are not only common to the real-world, but also relevant to target label combinations. Finally, ATOM combines metamorphic testing and label information to realize test oracle identification, based on which the ability of MICS in classifying different label combinations is evaluated. To evaluate the effectiveness of ATOM, we have performed experiments on two popular datasets of MICS, VOC and COCO (each with five state-of-the-art DNN models), and one real-world photo tagging application from our industrial partner. The experimental results reveal that the performance of current DNN-based MICSs remains less satisfactory even in recognizing correlations between only two labels, as ATOM triggers a total number of 6,049 such label combination related errors for all MICSs studied. In particular, ATOM reports 587 error-revealing images for the industrial MICS, in which 92% of them are confirmed by the developers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao 等ISSTA 2022 · 被引用 24 次
- Metamorphic Object Insertion for Testing Object Detection SystemsShuai Wang, Zhendong SuASE 2020 · 被引用 69 次
- Testing DNN image classifiers for confusion & bias errorsYuchi Tian, Ziyuan Zhong, Vicente Ordonez, Gail E. Kaiser 等ICSE 2020 · 被引用 33 次
- ASRTest: automated testing for deep-neural-network-driven speech recognition systemsPin Ji, Yang Feng, Jia Liu, Zhihong Zhao 等ISSTA 2022 · 被引用 22 次
- A Miss Is as Good as A Mile: Metamorphic Testing for Deep Learning OperatorsJinyin Chen, Chengyu Jia, Yunjie Yan, Jie Ge 等FSE 2024 · 被引用 8 次
