Multiple-Boundary Clustering and Prioritization to Promote Neural Network Retraining
Weijun Shen, Yanhui Li, Lin Chen, Yuanlei Han, Yuming Zhou, Baowen Xu
Abstract
With the increasing application of deep learning (DL) models in many safety-critical scenarios, effective and efficient DL testing techniques are much in demand to improve the quality of DL models. One of the major challenges is the data gap between the training data to construct the models and the testing data to evaluate them. To bridge the gap, testers aim to collect an effective subset of inputs from the testing contexts, with limited labeling effort, for retraining DL models.
To assist the subset selection, we propose Multiple-Boundary Clustering and Prioritization (MCP), a technique to cluster test samples into the boundary areas of multiple boundaries for DL models and specify the priority to select samples evenly from all boundary areas, to make sure enough useful samples for each boundary reconstruction. To evaluate MCP, we conduct an extensive empirical study with three popular DL models and 33 simulated testing contexts. The experiment results show that, compared with stateof-the-art baseline methods, on effectiveness, our approach MCP has a significantly better performance by evaluating the improved quality of retrained DL models; on efficiency, MCP also has the advantages in time costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 734ea284-bc22-4972-a8d2-a040c26498d2Cited by top-tier papers9
- TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning TasksYu Li, Min Li, Qiuxia Lai, Yannan Liu et al.NeurIPS 2021 · 36 citations
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- Measuring Discrimination to Boost Comparative Testing for Multiple Deep Learning ModelsLinghan Meng, Yanhui Li, Lin Chen, Zhi Wang et al.ICSE 2021 · 21 citations
- CertPri: Certifiable Prioritization for Deep Neural Networks via Movement Cost in Feature SpaceHaibin Zheng, Jinyin Chen, Haibo JinASE 2023 · 11 citations
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu et al.FSE 2023 · 10 citations
Builds on8
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- DeepBillboard: systematic physical-world testing of autonomous driving systemsHusheng Zhou, Wei Li, Zelun Kong, Junfeng Guo et al.ICSE 2020 · 150 citations
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 116 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
Related papers
- Test Selection for Deep Neural Networks using Meta-Models with Uncertainty MetricsDemet Demir, Aysu Betin Can, Elif SürerISSTA 2024 · 3 citations
- Prioritizing Test Inputs for DNNs Using Training DynamicsJian Shen, Zhong Li, Minxue Pan, Xuandong LiASE 2024 · 1 citation
- Cats Are Not Fish: Deep Learning Testing Calls for Out-Of-Distribution AwarenessDavid Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou et al.ASE 2020 · 56 citations
- Test Case Prioritization for DNNs via Neural Collapse InstabilityChunyu Liu, Mingyuan Li, Yang Li, Wenmin Li et al.ISSTA 2026
- Prioritizing Test Inputs for Deep Neural Networks via Mutation AnalysisZan Wang, Hanmo You, Junjie Chen, Yingyi Zhang et al.ICSE 2021 · 117 citations
