Energy-based Automated Model Evaluation
Ru Peng, Heming Zou, Haobo Wang, Yawen Zeng, Zenan Huang, Junbo Zhao
摘要
The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real world applications. The Automated Model Evaluation (AutoEval) shows an alternative to this traditional workflow, by forming a proximal prediction pipeline of the testing performance without the presence of ground-truth labels. Despite its recent successes, the AutoEval frameworks still suffer from an overconfidence issue, substantial storage and computational cost. In that regard, we propose a novel measure -- Meta-Distribution Energy (MDE) -- that allows the AutoEval framework to be both more efficient and effective. The core of the MDE is to establish a meta-distribution statistic, on the information (energy) associated with individual samples, then offer a smoother representation enabled by energy-based learning. We further provide our theoretical insights by connecting the MDE with the classification loss. We provide extensive experiments across modalities, datasets and different architectural backbones to validate MDE's validity, together with its superiority compared with prior approaches. We also prove MDE's versatility by showing its seamless integration with large-scale models, and easy adaption to learning scenarios with noisy- or imbalanced- labels. Code and data are available: https://github.com/pengr/Energy_AutoEval
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li 等WWW 2024 · 被引用 23 次
- Learning High-Order Relationships of Brain RegionsWeikang Qiu, Huangrui Chu, Selena Wang, Haolan Zuo 等ICML 2024 · 被引用 14 次
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang 等ICML 2026 · 被引用 13 次
- MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution ShiftsRenchunzi Xie, Ambroise Odonnat, Vasilii Feofanov, Weijian Deng 等NeurIPS 2024 · 被引用 10 次
- Towards Unsupervised Model Selection for Domain Adaptive Object DetectionHengfu Yu, Jinhong Deng, Wen Li, Lixin DuanNeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper35
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
相关 Paper
- CAME: Contrastive Automated Model EvaluationRu Peng, Qiuyang Duan, Haobo Wang, Jiachen Ma 等ICCV 2023 · 被引用 8 次
- Automated Model Evaluation for Object Detection Via Prediction Consistency and ReliabilitySeungju Yoo, Hyuk Kwon, Joong-Won Hwang, Kibok LeeICCV 2025 · 被引用 1 次
- Are Labels Always Necessary for Classifier Accuracy Evaluation?Weijian Deng, Liang ZhengCVPR 2021
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen 等KDD 2026 · 被引用 1 次
- Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution DetectionQi Chen, Hu DingCVPR 2025
