CAME: Contrastive Automated Model Evaluation
Ru Peng, Qiuyang Duan, Haobo Wang, Jiachen Ma, Yanbo Jiang, Yongjun Tu, Xiu Jiang, Junbo Zhao
摘要
The Automated Model Evaluation (AutoEval) framework entertains the possibility of evaluating a trained machine learning model without resorting to a labeled testing set. Despite the promise and some decent results, the existing AutoEval methods heavily rely on computing distribution shifts between the unlabelled testing set and the training set. We believe this reliance on the training set becomes another obstacle in shipping this technology to real-world ML development. In this work, we propose Contrastive Automatic Model Evaluation (CAME), a novel AutoEval framework that is rid of involving training set in the loop. The core idea of CAME bases on a theoretical analysis which bonds the model performance with a contrastive loss. Further, with extensive empirical validation, we manage to set up a predictable relationship between the two, simply by deducing on the unlabeled/unseen testing set. The resulting framework CAME establishes a new SOTA results for AutoEval by surpassing prior work significantly. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 被引用 10 次
- Towards Unsupervised Model Selection for Domain Adaptive Object DetectionHengfu Yu, Jinhong Deng, Wen Li, Lixin DuanNeurIPS 2024 · 被引用 7 次
- MuggleMath: Assessing the Impact of Query and Response Augmentation on Math ReasoningChengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong 等ACL 2024 · 被引用 4 次
- Automated Model Evaluation for Object Detection Via Prediction Consistency and ReliabilitySeungju Yoo, Hyuk Kwon, Joong-Won Hwang, Kibok LeeICCV 2025 · 被引用 1 次
它引用的顶会 Paper25
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Scaling Out-of-Distribution Detection for Real-World SettingsDan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou 等ICML 2022 · 被引用 653 次
- Model Adaptation: Historical Contrastive Learning for Unsupervised Domain Adaptation without Source DataJiaxing Huang, Dayan Guan, Aoran Xiao, Shijian LuNeurIPS 2021 · 被引用 301 次
相关 Paper
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng 等ICLR 2024 · 被引用 18 次
- Are Labels Always Necessary for Classifier Accuracy Evaluation?Weijian Deng, Liang ZhengCVPR 2021
- Universal Novelty Detection Through Adaptive Contrastive LearningHossein Mirzaei, Mojtaba Nafez, Mohammad Jafari, Mohammad Bagher Soltani 等CVPR 2024 · 被引用 9 次
- A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and OptimizationAnda Cheng, Wei Huang, Yinggui WangEMNLP 2025
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen 等KDD 2026 · 被引用 1 次
