Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled Data
Zhilin Zhao, Longbing Cao, Xuhui Fan, Wei-Shi Zheng
摘要
There are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorithm on training samples remains applicable to test samples. We present a novel approach called Importance Divergence (I-Div) to address the challenge of test label unavailability, enabling distribution discrepancy evaluation using only training samples. I-Div transfers the sampling patterns from the test distribution to the training distribution by estimating density and likelihood ratios. Specifically, the density ratio, informed by the selected hypothesis, is obtained by minimizing the Kullback-Leibler divergence between the actual and estimated input distributions. Simultaneously, the likelihood ratio is adjusted according to the density ratio by reducing the generalization error of the distribution discrepancy as transformed through the two ratios. Experimentally, I-Div accurately quantifies the distribution discrepancy, as evidenced by a wide range of complex data scenarios and tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang 等ICML 2020 · 被引用 213 次
相关 Paper
- R-divergence for Estimating Model-oriented Distribution DiscrepancyZhilin Zhao, Longbing CaoNeurIPS 2023 · 被引用 3 次
- KL Guided Domain AdaptationA. Tuan Nguyen, Toan Tran, Yarin Gal, Philip H. S. Torr 等ICLR 2022
- (Almost) Provable Error Bounds Under Distribution Shift via Disagreement DiscrepancyElan Rosenfeld, Saurabh GargNeurIPS 2023 · 被引用 18 次
- On the Importance of Feature Separability in Predicting Out-Of-Distribution ErrorRenchunzi Xie, Hongxin Wei, Lei Feng, Yuzhou Cao 等NeurIPS 2023 · 被引用 19 次
- f-Domain Adversarial Learning: Theory and AlgorithmsDavid Acuna, Guojun Zhang, Marc T. Law, Sanja FidlerICML 2021 · 被引用 77 次
