SAND: One-Shot Feature Selection with Additive Noise Distortion
Pedram Pad, Hadi Hammoud, Mohamad Dia, Nadim Maamari, Liza Andrea Dunbar
Abstract
Feature selection is a critical step in data-driven applications, reducing input dimensionality to enhance learning accuracy, computational efficiency, and interpretability. Existing state-of-theart methods often require post-selection retraining and extensive hyperparameter tuning, complicating their adoption. We introduce a novel, non-intrusive feature selection layer that, given a target feature count k, automatically identifies and selects the k most informative features during neural network training. Our method is uniquely simple, requiring no alterations to the loss function, network architecture, or post-selection retraining. The layer is mathematically elegant and can be fully described by: xi = a i x i + (1 -a i )z i where x i is the input feature, xi the output, z i a Gaussian noise, and a i trainable gain such that i a 2 i = k. This formulation induces an automatic clustering effect, driving k of the a i gains to 1 (selecting informative features) and the rest to 0 (discarding redundant ones) via weighted noise distortion and gain normalization. Despite its extreme simplicity, our method achieves competitive performance on standard benchmark datasets and a novel real-world dataset, often matching or exceeding existing approaches without requiring hyperparameter search for k or retraining. Theoretical analysis in the context of linear regression further validates its efficacy. Our work demonstrates that simplicity and performance are not mutually exclusive, offering a powerful yet straightforward tool for feature selection in machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9541b74b-e2f0-4e67-bf5a-7dcb34f876faBuilds on3
- Feature Selection using Stochastic GatesYutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval KlugerICML 2020 · 39 citations
- Where to Pay Attention in Sparse Training for Feature Selection?Ghada Sokar, Zahra Atashgahi, Mykola Pechenizkiy, Decebal Constantin MocanuNeurIPS 2022 · 25 citations
- Sequential Attention for Feature SelectionTaisuke Yasuda, Mohammad Hossein Bateni, Lin Chen, Matthew Fahrbach et al.ICLR 2023 · 2 citations
Related papers
- Overcoming Simplicity Bias in Deep Networks using a Feature SieveRishabh Tiwari, Pradeep ShenoyICML 2023 · 32 citations
- Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature EfficiencyNaoki Nishikawa, Rei Higuchi, Taiji SuzukiNeurIPS 2025 · 2 citations
- Fractal Autoencoders for Feature SelectionXinxing Wu, Qiang ChengAAAI 2021 · 33 citations
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li et al.ICLR 2020 · 83 citations
- NEAR: A Training-Free Pre-Estimator of Machine Learning Model PerformanceRaphael T. Husistein, Markus Reiher, Marco EckhoffICLR 2025
