Energy-based Automated Model Evaluation
Ru Peng, Heming Zou, Haobo Wang, Yawen Zeng, Zenan Huang, Junbo Zhao
Abstract
The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real world applications. The Automated Model Evaluation (AutoEval) shows an alternative to this traditional workflow, by forming a proximal prediction pipeline of the testing performance without the presence of ground-truth labels. Despite its recent successes, the AutoEval frameworks still suffer from an overconfidence issue, substantial storage and computational cost. In that regard, we propose a novel measure -- Meta-Distribution Energy (MDE) -- that allows the AutoEval framework to be both more efficient and effective. The core of the MDE is to establish a meta-distribution statistic, on the information (energy) associated with individual samples, then offer a smoother representation enabled by energy-based learning. We further provide our theoretical insights by connecting the MDE with the classification loss. We provide extensive experiments across modalities, datasets and different architectural backbones to validate MDE's validity, together with its superiority compared with prior approaches. We also prove MDE's versatility by showing its seamless integration with large-scale models, and easy adaption to learning scenarios with noisy- or imbalanced- labels. Code and data are available: https://github.com/pengr/Energy_AutoEval
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f77dbfb0-01c4-439f-9101-c9b8a06cfa38Cited by top-tier papers11
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li et al.WWW 2024 · 23 citations
- Learning High-Order Relationships of Brain RegionsWeikang Qiu, Huangrui Chu, Selena Wang, Haolan Zuo et al.ICML 2024 · 14 citations
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang et al.ICML 2026 · 13 citations
- MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution ShiftsRenchunzi Xie, Ambroise Odonnat, Vasilii Feofanov, Weijian Deng et al.NeurIPS 2024 · 10 citations
- Towards Unsupervised Model Selection for Domain Adaptive Object DetectionHengfu Yu, Jinhong Deng, Wen Li, Lixin DuanNeurIPS 2024 · 7 citations
Builds on35
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
Related papers
- CAME: Contrastive Automated Model EvaluationRu Peng, Qiuyang Duan, Haobo Wang, Jiachen Ma et al.ICCV 2023 · 8 citations
- Automated Model Evaluation for Object Detection Via Prediction Consistency and ReliabilitySeungju Yoo, Hyuk Kwon, Joong-Won Hwang, Kibok LeeICCV 2025 · 1 citation
- Are Labels Always Necessary for Classifier Accuracy Evaluation?Weijian Deng, Liang ZhengCVPR 2021
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
- Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution DetectionQi Chen, Hu DingCVPR 2025
