Characterizing Out-of-Distribution Error via Optimal Transport
Yuzhe Lu, Yilong Qin, Runtian Zhai, Andrew Shen, Ketong Chen, Zhenlin Wang, Soheil Kolouri, Simon Stepputtis, Joseph Campbell, Katia P. Sycara
摘要
Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the actual error, sometimes by a large margin, which greatly impacts their applicability to real tasks. In this work, we identify pseudo-label shift, or the difference between the predicted and true OOD label distributions, as a key indicator to this under-estimation. Based on this observation, we introduce a novel method for estimating model performance by leveraging optimal transport theory, Confidence Optimal Transport (COT), and show that it provably provides more robust error estimates in the presence of pseudo label shift. Additionally, we introduce an empirically-motivated variant of COT, Confidence Optimal Transport with Thresholding (COTT), which applies thresholding to the individual transport costs and further improves the accuracy of COT's error estimates. We evaluate COT and COTT on a variety of standard benchmarks that induce various types of distribution shift -synthetic, novel subpopulation, and natural -and show that our approaches significantly outperform existing state-of-the-art methods with up to 3x lower prediction errors. Our code can be found at https://github.com/luyuzhe111/COT . Performance prediction on unlabeled data has previously been shown to be impossible without imposing additional constraints over the unknown target distribution [7, 11, 5, 26] , due to the fact that target samples may take any label. Thus, the feasibility of this task is dependent on what assumptions we make regarding the shift between the train and target distributions. Prior works often make the assumption that the conditional density P (y|x) remains fixed in the presence of covariate shift [37] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng 等ICLR 2024 · 被引用 18 次
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 被引用 10 次
- MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution ShiftsRenchunzi Xie, Ambroise Odonnat, Vasilii Feofanov, Weijian Deng 等NeurIPS 2024 · 被引用 10 次
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag 等NeurIPS 2025 · 被引用 9 次
- Towards Unsupervised Model Selection for Domain Adaptive Object DetectionHengfu Yu, Jinhong Deng, Wen Li, Lixin DuanNeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper14
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 被引用 267 次
- Language-Conditioned Imitation Learning for Robot Manipulation TasksSimon Stepputtis, Joseph Campbell, Mariano J. Phielipp, Stefan Lee 等NeurIPS 2020 · 被引用 258 次
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 被引用 193 次
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur 等ICLR 2022 · 被引用 160 次
相关 Paper
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled DataAlvaro H. C. Correia, Christos LouizosNeurIPS 2025 · 被引用 5 次
- ODP-Bench: Benchmarking Out-Of-Distribution Performance PredictionHan Yu, Kehan Li, Dongbai Li, Yue He 等ICCV 2025
- Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal TransportYang Xiao, Weiming Liu, Jun Dan, Tengyue Xu 等CVPR 2026
