Feature Selection using Stochastic Gates
Yutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval Kluger
Abstract
Feature selection problems have been extensively studied in the setting of linear estimation (e.g. LASSO), but less emphasis has been placed on feature selection for non-linear functions. In this study, we propose a method for feature selection in neural network estimation problems. The new procedure is based on probabilistic relaxation of the 0 norm of features, or the count of the number of selected features. Our 0 -based regularization relies on a continuous relaxation of the Bernoulli distribution; such relaxation allows our model to learn the parameters of the approximate Bernoulli distributions via gradient descent. The proposed framework simultaneously learns either a nonlinear regression or classification function while selecting a small subset of features. We provide an information-theoretic justification for incorporating Bernoulli distribution into feature selection. Furthermore, we evaluate our method using synthetic and real-life data to demonstrate that our approach outperforms other commonly used methods in both predictive performance and feature selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f60546de-c6f6-4493-a4ea-e9f1653ebe71Cited by top-tier papers50
- Learning to Maximize Mutual Information for Dynamic Feature SelectionIan Connick Covert, Wei Qiu, Mingyu Lu, Nayoon Kim et al.ICML 2023 · 67 citations
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde et al.NeurIPS 2024 · 57 citations
- FedSDG-FS: Efficient and Secure Feature Selection for Vertical Federated LearningAnran Li, Hongyi Peng, Lan Zhang, Jiahui Huang et al.INFOCOM 2023 · 50 citations
- Locally Sparse Neural Networks for Tabular Biomedical DataJunchen Yang, Ofir Lindenbaum, Yuval KlugerICML 2022 · 45 citations
- Differentiable Unsupervised Feature Selection based on a Gated LaplacianOfir Lindenbaum, Uri Shaham, Erez Peterfreund, Jonathan Svirsky et al.NeurIPS 2021 · 38 citations
Related papers
- Discovering Features with Synergistic Interactions in Multiple ViewsChohee Kim, Mihaela van der Schaar, Changhee LeeICML 2024 · 4 citations
- Few-shot Learning for Feature Selection with Hilbert-Schmidt Independence CriterionAtsutoshi Kumagai, Tomoharu Iwata, Yasutoshi Ida, Yasuhiro FujiwaraNeurIPS 2022 · 13 citations
- Training Binary Neural Networks using the Bayesian Learning RuleXiangming Meng, Roman Bachmann, Mohammad Emtiyaz KhanICML 2020 · 47 citations
- Safe screening rules for L0-regression from Perspective RelaxationsAlper Atamtürk, Andrés GómezICML 2020 · 12 citations
- Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?Kaiqi Zhang, Yu-Xiang WangICLR 2023 · 3 citations
