Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing Approach
Zhijian Zhou, Liuhua Peng, Xunye Tian, Mingming Gong, Feng Liu
摘要
Are two distributions close to each other with statistical significance? Distribution closeness testing (DCT) formalizes this question by testing whether the distance between a distribution pair is at least -far. Existing DCT methods mainly measure discrepancies between a distribution pair defined on discrete spaces (e.g., using total variation), which limits their applications to complex data (e.g., images). To extend DCT to more types of data, a natural idea is to introduce maximum mean discrepancy (MMD), a powerful measurement of the distributional discrepancy between two complex distributions, into DCT scenarios. However, the empirical results indicate that many distribution pairs can have the same MMD value despite having different norms in the same reproducing kernel Hilbert space (RKHS), and these pairs may exhibit different finite-sample distinguishability and reflect different practical closeness levels, making MMD less informative in DCT. To mitigate the issue, we design a new measurement of distributional discrepancy, norm-adaptive MMD (NAMMD), which scales MMD's value using the RKHS norms of distributions. Based on the asymptotic distribution of NAMMD, we finally propose the NAMMD-based DCT to assess the closeness level of a distribution pair. Theoretically, we prove that NAMMD-based DCT has higher test power compared to MMD-based DCT, with bounded type-I error, which is also validated by extensive experiments on many types of data (e.g., synthetic noise, real images). Our code is available at: https://github.com/zhijianzhouml/NAMMD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang 等ICML 2020 · 被引用 213 次
- Is Out-of-Distribution Detection Learnable?Zhen Fang, Yixuan Li, Jie Lu, Jiahua Dong 等NeurIPS 2022 · 被引用 188 次
- Maximum Mean Discrepancy Test is Aware of Adversarial AttacksRuize Gao, Feng Liu, Jingfeng Zhang, Bo Han 等ICML 2021 · 被引用 77 次
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 被引用 49 次
- Higher Order Kernel Mean Embeddings to Capture Filtrations of Stochastic ProcessesCristopher Salvi, Maud Lemercier, Chong Liu, Blanka Horvath 等NeurIPS 2021 · 被引用 43 次
相关 Paper
- Neural Tangent Kernel Maximum Mean DiscrepancyXiuyuan Cheng, Yao XieNeurIPS 2021 · 被引用 26 次
- Kernel-based Maximum-of-difference Test for Two-sample ComparisonDan Pu, Tianyi Zhu, Yao Yan, Wei LanICML 2026
- A permutation-free kernel two-sample testShubhanshu Shekhar, Ilmun Kim, Aaditya RamdasNeurIPS 2022 · 被引用 40 次
- Anchor-based Maximum Discrepancy for Relative Similarity TestingZhijian Zhou, Liuhua Peng, Xunye Tian, Feng LiuNeurIPS 2025 · 被引用 2 次
- Kernel Quantile Embeddings and Associated Probability MetricsMasha Naslidnyk, Siu Lun Chau, François-Xavier Briol, Krikamol MuandetICML 2025
