Consistent feature selection for analytic deep neural networks
Vu C. Dinh, Lam Si Tung Ho
Abstract
One of the most important steps toward interpretability and explainability of neural network models is feature selection, which aims to identify the subset of relevant features. Theoretical results in the field have mostly focused on the prediction aspect of the problem with virtually no work on feature selection consistency for deep neural networks due to the model's severe nonlinearity and unidentifiability. This lack of theoretical foundation casts doubt on the applicability of deep learning to contexts where correct interpretations of the features play a central role. In this work, we investigate the problem of feature selection for analytic deep networks. We prove that for a wide class of networks, including deep feed-forward neural networks, convolutional neural networks, and a major sub-class of residual neural networks, the Adaptive Group Lasso selection procedure with Group Lasso as the base estimator is selection-consistent. The work provides further evidence that Group Lasso might be inefficient for feature selection with neural networks and advocates the use of Adaptive Group Lasso over the popular Group Lasso.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65c8bf38-a368-4f6a-b4e7-0438277b180fCited by top-tier papers10
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang et al.CVPR 2022 · 149 citations
- Neural graphical modelling in continuous-time: consistency guarantees and algorithmsAlexis Bellot, Kim Branson, Mihaela van der SchaarICLR 2022 · 57 citations
- LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated LearningTimothy Castiglia, Yi Zhou, Shiqiang Wang, Swanand Kadhe et al.ICML 2023 · 33 citations
- Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Sparse Neural NetworksShuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen et al.NeurIPS 2021 · 28 citations
- Bilevel Network Learning via Hierarchically Structured SparsityJiayi Fan, Jingyuan Yang, Shuangge Ma, Mengyun WuNeurIPS 2025 · 1 citation
Related papers
- Global Optimality Beyond Two Layers: Training Deep ReLU Networks via Convex ProgramsTolga Ergen, Mert PilanciICML 2021 · 35 citations
- The Contextual Lasso: Sparse Linear Models via Deep Neural NetworksRyan Thompson, Amir Dezfouli, Robert KohnNeurIPS 2023 · 8 citations
- Algorithmic stability and generalization of an unsupervised feature selection algorithmXinxing Wu, Qiang ChengNeurIPS 2021 · 13 citations
- Theoretical Characterisation of the Gauss Newton Conditioning in Neural NetworksJim Zhao, Sidak Pal Singh, Aurélien LucchiNeurIPS 2024 · 7 citations
- Architectural Adversarial Robustness: The Case for Deep PursuitGeorge Cazenavette, Calvin Murdock, Simon LuceyCVPR 2021
