(Almost) Provable Error Bounds Under Distribution Shift via Disagreement Discrepancy
Elan Rosenfeld, Saurabh Garg
摘要
We derive an (almost) guaranteed upper bound on the error of deep neural networks under distribution shift using unlabeled test data. Prior methods either give bounds that are vacuous in practice or give estimates that are accurate on average but heavily underestimate error for a sizeable fraction of shifts. In particular, the latter only give guarantees based on complex continuous measures such as test calibration -- which cannot be identified without labels -- and are therefore unreliable. Instead, our bound requires a simple, intuitive condition which is well justified by prior empirical works and holds in practice effectively 100% of the time. The bound is inspired by -divergence but is easier to evaluate and substantially tighter, consistently providing non-vacuous guarantees. Estimating the bound requires optimizing one multiclass classifier to disagree with another, for which some prior works have used sub-optimal proxy losses; we devise a"disagreement loss"which is theoretically justified and performs better in practice. We expect this loss can serve as a drop-in replacement for future methods which require maximizing multiclass disagreement. Across a wide range of benchmarks, our method gives valid error bounds while achieving average accuracy comparable to competitive estimation baselines. Code is publicly available at https://github.com/erosenfeld/disagree_discrep .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- When is Multicalibration Post-Processing Necessary?Dutch Hansen, Siddartha Devic, Preetum Nakkiran, Vatsal SharanNeurIPS 2024 · 被引用 21 次
- Monitoring Risks in Test-Time AdaptationMona Schirmer, Metod Jazbec, Christian Andersson Naesseth, Eric T. NalisnickNeurIPS 2025 · 被引用 10 次
- Reliably detecting model failures in deployment without labelsViet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya 等NeurIPS 2025 · 被引用 3 次
- Knowledge Distillation of Uncertainty using Deep Latent Factor ModelSehyun Park, Jongjin Lee, Yunseop Shin, Ilsang Ohn 等NeurIPS 2025 · 被引用 2 次
- Bridging Domain Expertise and Generalization for Performance EstimationShuxuan Li, Zhilin Zhao, Quyu Kong, Wei-Shi ZhengCVPR 2026 · 被引用 1 次
它引用的顶会 Paper23
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
相关 Paper
- On the Bayes Inconsistency of Disagreement Discrepancy SurrogatesNeil G Marchant, Andrew Craig Cullen, Feng Liu, Sarah Monazam ErfaniICLR 2026
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 被引用 134 次
- Bayesian Adaptation for Covariate ShiftAurick Zhou, Sergey LevineNeurIPS 2021 · 被引用 40 次
- RATT: Leveraging Unlabeled Data to Guarantee GeneralizationSaurabh Garg, Sivaraman Balakrishnan, J. Zico Kolter, Zachary C. LiptonICML 2021 · 被引用 30 次
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
