Are Labels Always Necessary for Classifier Accuracy Evaluation?
Weijian Deng, Liang Zheng
Abstract
To calculate the model accuracy on a computer vision task, e.g., object recognition, we usually require a test set composing of test samples and their ground truth labels. Whilst standard usage cases satisfy this requirement, many real-world scenarios involve unlabeled test data, rendering common model evaluation methods infeasible. We investigate this important and under-explored problem, Automatic model Evaluation (AutoEval). Specifically, given a labeled training set and a classifier, we aim to estimate the classification accuracy on unlabeled test datasets. We construct a meta-dataset: a dataset comprised of datasets generated from the original images via various transformations such as rotation, background substitution, foreground scaling, etc. As the classification accuracy of the model on each sample (dataset) is known from the original dataset labels, our task can be solved via regression. Using the feature statistics to represent the distribution of a sample dataset, we can train regression models (e.g., a regression neural network) to predict model performance. Using synthetic meta-dataset and real-world datasets in training and testing, respectively, we report a reasonable and promising prediction of the model accuracy. We also provide insights into the application scope, limitation, and potential future direction of AutoEval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1ce8c12-1c98-4151-a5da-c80527af2e0bCited by top-tier papers51
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu et al.NeurIPS 2021 · 79 citations
- What Does Rotation Prediction Tell Us about Classifier Accuracy under Varying Testing Environments?Weijian Deng, Stephen Gould, Liang ZhengICML 2021 · 72 citations
Builds on4
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Computing the Testing Error Without a Testing SetCiprian A. Corneanu, Sergio Escalera, Aleix M. MartinezCVPR 2020
Related papers
- On the Evaluation of Capability Estimation Methods for Large Language ModelsQiang Hu, Jin Wen, Yao Zhang, Maxime Cordy et al.AAAI 2026
- Automated Model Evaluation for Object Detection Via Prediction Consistency and ReliabilitySeungju Yoo, Hyuk Kwon, Joong-Won Hwang, Kibok LeeICCV 2025 · 1 citation
- CAME: Contrastive Automated Model EvaluationRu Peng, Qiuyang Duan, Haobo Wang, Jiachen Ma et al.ICCV 2023 · 8 citations
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
