I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively
Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma
摘要
The learning of hierarchical representations for image classification has experienced an impressive series of successes due in part to the availability of large-scale labeled data for training. On the other hand, the trained classifiers have traditionally been evaluated on small and fixed sets of test images, which are deemed to be extremely sparsely distributed in the space of all natural images. It is thus questionable whether recent performance improvements on the excessively re-used test sets generalize to real-world natural images with much richer content variations. Inspired by efficient stimulus selection for testing perceptual models in psychophysical and physiological studies, we present an alternative framework for comparing image classifiers, which we name the MAximum Discrepancy (MAD) competition. Rather than comparing image classifiers using fixed test images, we adaptively sample a small test set from an arbitrarily large corpus of unlabeled images so as to maximize the discrepancies between the classifiers, measured by the distance over WordNet hierarchy. Human labeling on the resulting model-dependent image sets reveals the relative performance of the competing classifiers, and provides useful insights on potential ways to improve them. We report the MAD competition results of eleven ImageNet classifiers while noting that the framework is readily extensible and cost-effective to add future classifiers into the competition. Codes can be found at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- AugMax: Adversarial Composition of Random Augmentations for Robust TrainingHaotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu 等NeurIPS 2021 · 被引用 153 次
- Self-Damaging Contrastive LearningZiyu Jiang, Tianlong Chen, Bobak J. Mortazavi, Zhangyang WangICML 2021 · 被引用 83 次
- Adaptive Testing of Computer Vision ModelsIrena Gao, Gabriel Ilharco, Scott M. Lundberg, Marco Túlio RibeiroICCV 2023 · 被引用 49 次
- Learning Where to Edit Vision TransformersYunqiao Yang, Long-Kai Huang, Shengzhuang Chen, Kede Ma 等NeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
- Agreement-Discrepancy-Selection: Active Learning with Progressive Distribution AlignmentMengying Fu, Tianning Yuan, Fang Wan, Songcen Xu 等AAAI 2021 · 被引用 13 次
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 等ICLR 2020 · 被引用 11 次
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner 等CVPR 2024
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 被引用 103 次
