I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively
Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma
Abstract
The learning of hierarchical representations for image classification has experienced an impressive series of successes due in part to the availability of large-scale labeled data for training. On the other hand, the trained classifiers have traditionally been evaluated on small and fixed sets of test images, which are deemed to be extremely sparsely distributed in the space of all natural images. It is thus questionable whether recent performance improvements on the excessively re-used test sets generalize to real-world natural images with much richer content variations. Inspired by efficient stimulus selection for testing perceptual models in psychophysical and physiological studies, we present an alternative framework for comparing image classifiers, which we name the MAximum Discrepancy (MAD) competition. Rather than comparing image classifiers using fixed test images, we adaptively sample a small test set from an arbitrarily large corpus of unlabeled images so as to maximize the discrepancies between the classifiers, measured by the distance over WordNet hierarchy. Human labeling on the resulting model-dependent image sets reveals the relative performance of the competing classifiers, and provides useful insights on potential ways to improve them. We report the MAD competition results of eleven ImageNet classifiers while noting that the framework is readily extensible and cost-effective to add future classifiers into the competition. Codes can be found at this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- AugMax: Adversarial Composition of Random Augmentations for Robust TrainingHaotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu et al.NeurIPS 2021 · 153 citations
- Self-Damaging Contrastive LearningZiyu Jiang, Tianlong Chen, Bobak J. Mortazavi, Zhangyang WangICML 2021 · 83 citations
- Adaptive Testing of Computer Vision ModelsIrena Gao, Gabriel Ilharco, Scott M. Lundberg, Marco Túlio RibeiroICCV 2023 · 49 citations
- Learning Where to Edit Vision TransformersYunqiao Yang, Long-Kai Huang, Shengzhuang Chen, Kede Ma et al.NeurIPS 2024 · 6 citations
Builds on1
Related papers
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma et al.ACL 2025
- Agreement-Discrepancy-Selection: Active Learning with Progressive Distribution AlignmentMengying Fu, Tianning Yuan, Fang Wan, Songcen Xu et al.AAAI 2021 · 13 citations
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner et al.CVPR 2024
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 103 citations
