Zipper: Addressing Degeneracy in Algorithm-Agnostic Inference
Geng Chen, Yinxu Jia, Guanghui Wang, Changliang Zou
Abstract
The widespread use of black box prediction methods has sparked an increasing interest in algorithm/model-agnostic approaches for quantifying goodness-of-fit, with direct ties to specification testing, model selection and variable importance assessment. A commonly used framework involves defining a predictiveness criterion, applying a cross-fitting procedure to estimate the predictiveness, and utilizing the difference in estimated predictiveness between two models as the test statistic. However, even after standardization, the test statistic typically fails to converge to a non-degenerate distribution under the null hypothesis of equal goodness, leading to what is known as the degeneracy issue. To addresses this degeneracy issue, we present a simple yet effective device, Zipper. It draws inspiration from the strategy of additional splitting of testing data, but encourages an overlap between two testing data splits in predictiveness evaluation. Zipper binds together the two overlapping splits using a slider parameter that controls the proportion of overlap. Our proposed test statistic follows an asymptotically normal distribution under the null hypothesis for any fixed slider value, guaranteeing valid size control while enhancing power by effective data reuse. Finite-sample experiments demonstrate that our procedure, with a simple choice of the slider, works well across a wide range of settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
- Cross-validation Confidence Intervals for Test ErrorPierre Bayle, Alexandre Bayle, Lucas Janson, Lester MackeyNeurIPS 2020 · 76 citations
- Lazy Estimation of Variable Importance for Large Neural NetworksYue Gao, Abby Stevens, Garvesh Raskutti, Rebecca WillettICML 2022 · 7 citations
Related papers
- Leveraging Predictive Equivalence in Decision TreesHayden McTavish, Zachery Boner, Jon Donnelly, Margo I. Seltzer et al.ICML 2025
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
- A Zest of LIME: Towards Architecture-Independent Model DistancesHengrui Jia, Hongyu Chen, Jonas Guan, Ali Shahin Shamsabadi et al.ICLR 2022 · 30 citations
- Conformal Prediction for Ensembles: Improving Efficiency via Score-Based AggregationYash Patel, Eduardo Ochoa Rivera, Ambuj TewariNeurIPS 2025 · 9 citations
- KSD Aggregated Goodness-of-fit TestAntonin Schrab, Benjamin Guedj, Arthur GrettonNeurIPS 2022 · 26 citations
