Aries: Efficient Testing of Deep Neural Networks via Labeling-Free Accuracy Estimation
Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Lei Ma, Yves Le Traon
Abstract
Deep learning (DL) plays a more and more important role in our daily life due to its competitive performance in industrial application domains. As the core of DL-enabled systems, deep neural networks (DNNs) need to be carefully evaluated to ensure the produced models match the expected requirements. In practice, the de facto standard to assess the quality of DNNs in the industry is to check their performance (accuracy) on a collected set of labeled test data. However, preparing such labeled data is often not easy partly because of the huge labeling effort, i.e., data labeling is labor-intensive, especially with the massive new incoming unlabeled data every day. Recent studies show that test selection for DNN is a promising direction that tackles this issue by selecting minimal representative data to label and using these data to assess the model. However, it still requires human effort and cannot be automatic. In this paper, we propose a novel technique, named Aries, that can estimate the performance of DNNs on new unlabeled data using only the information obtained from the original test data. The key insight behind our technique is that the model should have similar prediction accuracy on the data which have similar distances to the decision boundary. We performed a large-scale evaluation of our technique on two famous datasets, CIFAR-10 and Tiny-ImageNet, four widely studied DNN models including ResNetl0l and DenseNetl21, and 13 types of data transformation methods. Results show that the estimated accuracy by Aries is only 0.03% - 2.60% off the true accuracy. Besides, Aries also outperforms the state-of-the-art labeling-free methods in 50 out of 52 cases and selection-labeling-based methods in 96 out of 128 cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng et al.ICLR 2024 · 18 citations
- DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsLongtian Wang, Xiaofei Xie, Xiaoning Du, Meng Tian et al.FSE 2023 · 15 citations
- VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic ManipulationZhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang et al.FSE 2025 · 6 citations
- Decomposition of Deep Neural Networks into Modules via Mutation AnalysisAli GhanbariISSTA 2024 · 4 citations
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation TestingAli Ghanbari, Sasan TavakkolASE 2025 · 1 citation
Builds on13
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen et al.NeurIPS 2020 · 225 citations
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- Learnable Boundary Guided Adversarial TrainingJiequan Cui, Shu Liu, Liwei Wang, Jiaya JiaICCV 2021 · 152 citations
Related papers
- Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2021 · 3 citations
- Distance-Aware Test Input Selection for Deep Neural NetworksZhong Li, Zhengfeng Xu, Ruihua Ji, Minxue Pan et al.ISSTA 2024 · 4 citations
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Thomas RainforthNeurIPS 2022 · 36 citations
- Towards Exploring the Limitations of Active Learning: An Empirical StudyQiang Hu, Yuejun Guo, Maxime Cordy, Xiaofei Xie et al.ASE 2021 · 18 citations
- Test Selection for Deep Neural Networks using Meta-Models with Uncertainty MetricsDemet Demir, Aysu Betin Can, Elif SürerISSTA 2024 · 3 citations
