Statistical Test for Feature Selection Pipelines by Selective Inference
Tomohiro Shiraishi, Tatsuya Matsukawa, Shuichi Nishino, Ichiro Takeuchi
摘要
A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating various analysis algorithms. In this paper, we propose a novel statistical test to assess the significance of data analysis pipelines. Our approach enables the systematic development of valid statistical tests applicable to any feature selection pipeline composed of predefined components. We develop this framework based on selective inference, a statistical technique that has recently gained attention for data-driven hypotheses. As a proof of concept, we focus on feature selection pipelines for linear models, composed of three missing value imputation algorithms, three outlier detection algorithms, and three feature selection algorithms. We theoretically prove that our statistical test can control the probability of false positive feature selection at any desired level, and demonstrate its validity and effectiveness through experiments on synthetic and real data. Additionally, we present an implementation framework that facilitates testing across any configuration of these feature selection pipelines without extra implementation costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Computing Valid p-value for Optimal Changepoint by Selective Inference using Dynamic ProgrammingVo Nguyen Le Duy, Hiroki Toda, Ryota Sugiyama, Ichiro TakeuchiNeurIPS 2020 · 被引用 45 次
- Fast and More Powerful Selective Inference for Sparse High-Order Interaction ModelDiptesh Das, Vo Nguyen Le Duy, Hiroyuki Hanada, Koji Tsuda 等AAAI 2022 · 被引用 23 次
- Quantifying Statistical Significance of Neural Network-based Image Segmentation by Selective InferenceVo Nguyen Le Duy, Shogo Iwazaki, Ichiro TakeuchiNeurIPS 2022 · 被引用 21 次
- Computing Valid P-Values for Image Segmentation by Selective InferenceKosuke Tanizaki, Noriaki Hashimoto, Yu Inatsu, Hidekata Hontani 等CVPR 2020
相关 Paper
- More Powerful and General Selective Inference for Stepwise Feature Selection using Homotopy MethodKazuya Sugiyama, Vo Nguyen Le Duy, Ichiro TakeuchiICML 2021 · 被引用 18 次
- The Fault in our StatsAlexi Turcotte, Neev Nirav MehtaASE 2025
- Quantifying Statistical Significance of Deep Nearest Neighbor Anomaly Detection via Selective InferenceMizuki Niihori, Shuichi Nishino, Teruyuki Katsuoka, Tomohiro Shiraishi 等NeurIPS 2025 · 被引用 3 次
- Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter 等ICLR 2024 · 被引用 11 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
