How Does Unlabeled Data Provably Help Out-of-Distribution Detection?
Xuefeng Du, Zhen Fang, Ilias Diakonikolas, Yixuan Li
Abstract
Using unlabeled data to regularize the machine learning models has demonstrated promise for improving safety and reliability in detecting out-of-distribution (OOD) data. Harnessing the power of unlabeled in-the-wild data is non-trivial due to the heterogeneity of both in-distribution (ID) and OOD data. This lack of a clean set of OOD samples poses significant challenges in learning an optimal OOD classifier. Currently, there is a lack of research on formally understanding how unlabeled data helps OOD detection. This paper bridges the gap by introducing a new learning framework SAL (Separate And Learn) that offers both strong theoretical guarantees and empirical effectiveness. The framework separates candidate outliers from the unlabeled data and then trains an OOD classifier using the candidate outliers and the labeled ID data. Theoretically, we provide rigorous error bounds from the lens of separability and learnability, formally justifying the two components in our algorithm. Our theory shows that SAL can separate the candidate outliers with small error rates, which leads to a generalization guarantee for the learned OOD classifier. Empirically, SAL achieves state-of-the-art performance on common benchmarks, reinforcing our theoretical insights. Code is publicly available at https://github.com/deeplearning-wisc/sal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10083bde-126f-4ce9-8e8b-ab0bdb993e8eCited by top-tier papers13
- HaloScope: Harnessing Unlabeled LLM Generations for Hallucination DetectionXuefeng Du, Chaowei Xiao, Sharon LiNeurIPS 2024 · 131 citations
- GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution DetectionXin Gao, Jiyao Liu, Guanghao Li, Yueming Lyu et al.NeurIPS 2025 · 9 citations
- Bridging OOD Detection and Generalization: A Graph-Theoretic ViewHan Wang, Sharon LiNeurIPS 2024 · 7 citations
- Gradient Short-Circuit: Efficient Out-of-Distribution Detection via Feature InterventionJiawei Gu, Ziyue Qiao, Zechao LiICCV 2025 · 3 citations
- FEVER-OOD: Free Energy Vulnerability Elimination for Robust Out-of-Distribution DetectionBrian K. S. Isaac-Medina, Mauricio Che, Yona Falinie Binti A. Gaus, Samet Akcay et al.ICCV 2025 · 2 citations
Builds on40
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- ReAct: Out-of-distribution Detection With Rectified ActivationsYiyou Sun, Chuan Guo, Yixuan LiNeurIPS 2021 · 733 citations
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
Related papers
- Training OOD Detectors in their Natural HabitatsJulian Katz-Samuels, Julia B. Nakhleh, Robert D. Nowak, Yixuan LiICML 2022 · 115 citations
- When and How Does In-Distribution Label Help Out-of-Distribution Detection?Xuefeng Du, Yiyou Sun, Yixuan LiICML 2024 · 11 citations
- STEP: Out-of-Distribution Detection in the Presence of Limited In-Distribution Labeled DataZhi Zhou, Lan-Zhe Guo, Zhanzhan Cheng, Yufeng Li et al.NeurIPS 2021 · 41 citations
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon et al.ICML 2023 · 67 citations
- FedOpenMatch: Towards Semi-Supervised Federated Learning in Open-Set EnvironmentsHongquan Liu, ChenyuGuo Guo, Yixin Ren, Jihong Guan et al.ICLR 2026
