DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing
Zhijian Zhou, Xunye Tian, Liuhua Peng, Chao Lei, Antonin Schrab, Danica J. Sutherland, Feng Liu
Abstract
To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly maximizing multiple kernel-based statistics may result in highly similar kernels that capture highly overlapping information, limiting the effectiveness of aggregation. To address this, we propose an aggregated statistic that explicitly incorporates kernel diversity based on the covariance between different kernels. Moreover, we identify a fundamental challenge: a trade-off between the diversity among kernels and the test power of individual kernels, i.e., the selected kernels should be both effective and diverse. This motivates a testing framework with selection inference, which leverages information from the training phase to select kernels with strong individual performance from the learned diverse kernel pool. We provide rigorous theoretical statements and proofs to show the consistency on the test power and control of Type-I error, along with asymptotic analysis of the proposed statistics. Lastly, we conducted extensive empirical experiments demonstrating the superior performance of our proposed approach across various benchmarks for both twosample and independence testing. ¶
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 732c62de-379f-467e-b446-80d7eb7f5afdCited by top-tier papers3
- Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing ApproachZhijian Zhou, Liuhua Peng, Xunye Tian, Mingming Gong et al.ICML 2026 · 1 citation
- LOTTERY: Learning from Reference-Only Samples in Two-Sample Testing under Size AsymmetryXunye Tian, Zhijian Zhou, Liuhua Peng, Feng LiuICML 2026
- Adaptive Multiscale Binary Expansion Tests for IndependenceYang Yang, Duo Zheng, Sandeep Jain, Kai Zhang et al.ICML 2026
Builds on19
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu et al.NeurIPS 2024 · 124 citations
- Self-Supervised Learning with Kernel Dependence MaximizationYazhe Li, Roman Pogodin, Danica J. Sutherland, Arthur GrettonNeurIPS 2021 · 107 citations
- Maximum Mean Discrepancy Test is Aware of Adversarial AttacksRuize Gao, Feng Liu, Jingfeng Zhang, Bo Han et al.ICML 2021 · 77 citations
- Deep Unlearning via Randomized Conditionally Independent HessiansRonak Mehta, Sourav Pal, Vikas Singh, Sathya N. RaviCVPR 2022 · 50 citations
Related papers
- Practical Kernel Selection for Kernel-based Conditional Independence TestWenjie Wang, Mingming Gong, Biwei Huang, James Bailey et al.NeurIPS 2025 · 2 citations
- Meta Two-Sample Testing: Learning Kernels for Testing with Limited DataFeng Liu, Wenkai Xu, Jie Lu, Danica J. SutherlandNeurIPS 2021 · 30 citations
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 49 citations
- Kernel-based Maximum-of-difference Test for Two-sample ComparisonDan Pu, Tianyi Zhu, Yao Yan, Wei LanICML 2026
- Diversity-Aware Recursive Feature Multiple Kernel Learningnan cao, Xu Zhao, Teng ZhangICML 2026
