ODP-Bench: Benchmarking Out-Of-Distribution Performance Prediction
Han Yu, Kehan Li, Dongbai Li, Yue He, Xingxuan Zhang, Peng Cui
摘要
Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios. Although progress has been made in this area, evaluation protocols in previous literature are inconsistent, and most works cover only a limited number of real-world OOD datasets and types of distribution shifts. To provide convenient and fair comparisons for various algorithms, we propose Out-of-Distribution Performance Prediction Benchmark (ODP-Bench), a comprehensive benchmark that includes most commonly used OOD datasets and existing practical performance prediction algorithms. We provide our trained models as a testbench for future researchers, thus guaranteeing the consistency of comparison and avoiding the burden of repeating the model training process. Furthermore, we also conduct in-depth experimental analyses to better understand their capability boundary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Error Slice Discovery via Manifold CompactnessHan Yu, Hao Zou, Jiashuo Liu, Renzhe Xu 等AAAI 2026 · 被引用 2 次
- Generating Risky Samples with Conformity Constraints via Diffusion ModelsHan Yu, Hao Zou, Xingxuan Zhang, Zhengyi Wang 等AAAI 2026
它引用的顶会 Paper39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 被引用 2,652 次
相关 Paper
- OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution ShiftLin Li, Yifei Wang, Chawin Sitawarin, Michael W. SpratlingICML 2024 · 被引用 13 次
- RetroOOD: Understanding Out-of-Distribution Generalization in Retrosynthesis PredictionYemin Yu, Luotian Yuan, Ying Wei, Hanyu Gao 等AAAI 2024 · 被引用 5 次
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationNanyang Ye, Kaican Li, Haoyue Bai, Runpeng Yu 等CVPR 2022 · 被引用 74 次
- OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution DetectionMax Gutbrod, David Rauber, Danilo Weber Nunes, Christoph PalmCVPR 2025
- A framework for benchmarking Class-out-of-distribution detection and its application to ImageNetIdo Galil, Mohammed Dabbah, Ran El-YanivICLR 2023 · 被引用 2 次
