ODP-Bench: Benchmarking Out-Of-Distribution Performance Prediction
Han Yu, Kehan Li, Dongbai Li, Yue He, Xingxuan Zhang, Peng Cui
Abstract
Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios. Although progress has been made in this area, evaluation protocols in previous literature are inconsistent, and most works cover only a limited number of real-world OOD datasets and types of distribution shifts. To provide convenient and fair comparisons for various algorithms, we propose Out-of-Distribution Performance Prediction Benchmark (ODP-Bench), a comprehensive benchmark that includes most commonly used OOD datasets and existing practical performance prediction algorithms. We provide our trained models as a testbench for future researchers, thus guaranteeing the consistency of comparison and avoiding the burden of repeating the model training process. Furthermore, we also conduct in-depth experimental analyses to better understand their capability boundary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb4dde65-5196-48c3-80e3-c9fc1c2555c9Cited by top-tier papers2
- Error Slice Discovery via Manifold CompactnessHan Yu, Hao Zou, Jiashuo Liu, Renzhe Xu et al.AAAI 2026 · 2 citations
- Generating Risky Samples with Conformity Constraints via Diffusion ModelsHan Yu, Hao Zou, Xingxuan Zhang, Zhengyi Wang et al.AAAI 2026
Builds on39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 2,652 citations
Related papers
- OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution ShiftLin Li, Yifei Wang, Chawin Sitawarin, Michael W. SpratlingICML 2024 · 13 citations
- RetroOOD: Understanding Out-of-Distribution Generalization in Retrosynthesis PredictionYemin Yu, Luotian Yuan, Ying Wei, Hanyu Gao et al.AAAI 2024 · 5 citations
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationNanyang Ye, Kaican Li, Haoyue Bai, Runpeng Yu et al.CVPR 2022 · 74 citations
- OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution DetectionMax Gutbrod, David Rauber, Danilo Weber Nunes, Christoph PalmCVPR 2025
- A framework for benchmarking Class-out-of-distribution detection and its application to ImageNetIdo Galil, Mohammed Dabbah, Ran El-YanivICLR 2023 · 2 citations
