Measuring Discrimination to Boost Comparative Testing for Multiple Deep Learning Models
Linghan Meng, Yanhui Li, Lin Chen, Zhi Wang, Di Wu, Yuming Zhou, Baowen Xu
Abstract
The boom of DL technology leads to massive DL models built and shared, which facilitates the acquisition and reuse of DL models. For a given task, we encounter multiple DL models available with the same functionality, which are considered as candidates to achieve this task. Testers are expected to compare multiple DL models and select the more suitable ones w.r.t. the whole testing context. Due to the limitation of labeling effort, testers aim to select an efficient subset of samples to make an as precise rank estimation as possible for these models. To tackle this problem, we propose Sample Discrimination based Selection (SDS) to select efficient samples that could discriminate multiple models, i.e., the prediction behaviors (right/wrong) of these samples would be helpful to indicate the trend of model performance. To evaluate SDS, we conduct an extensive empirical study with three widely-used image datasets and 80 real world DL models. The experimental results show that, compared with state-of-the-art baseline methods, SDS is an effective and efficient sample selection method to rank multiple DL models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ccb1e4d-8452-4ca6-9a91-b719ad012cbeCited by top-tier papers3
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
- Reusing Deep Neural Network Models through Model Re-engineeringBinhang Qi, Hailong Sun, Xiang Gao, Hongyu Zhang et al.ICSE 2023 · 16 citations
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao et al.ICSE 2024 · 3 citations
Builds on7
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- Model-Reuse Attacks on Deep Learning SystemsYujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo et al.CCS 2018 · 197 citations
- DeepBillboard: systematic physical-world testing of autonomous driving systemsHusheng Zhou, Wei Li, Zelun Kong, Junfeng Guo et al.ICSE 2020 · 150 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 65 citations
Related papers
- Test Selection for Deep Neural Networks using Meta-Models with Uncertainty MetricsDemet Demir, Aysu Betin Can, Elif SürerISSTA 2024 · 3 citations
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 6 citations
- Multiple-Boundary Clustering and Prioritization to Promote Neural Network RetrainingWeijun Shen, Yanhui Li, Lin Chen, Yuanlei Han et al.ASE 2020 · 51 citations
- DISCO: Diversifying Sample Condensation for Efficient Model EvaluationAlexander Rubinstein, Benjamin Raible, Martin Gubri, Seong Joon OhICLR 2026 · 4 citations
- Towards Exploring the Limitations of Active Learning: An Empirical StudyQiang Hu, Yuejun Guo, Maxime Cordy, Xiaofei Xie et al.ASE 2021 · 18 citations
