Sequential Attention for Feature Selection
Taisuke Yasuda, Mohammad Hossein Bateni, Lin Chen, Matthew Fahrbach, Gang Fu, Vahab Mirrokni
Abstract
Feature selection is the problem of selecting a subset of features for a machine learning model that maximizes model quality subject to a budget constraint. For neural networks, prior methods, including those based on regularization, attention, and other techniques, typically select the entire feature subset in one evaluation round, ignoring the residual value of features during selection, i.e., the marginal contribution of a feature given that other features have already been selected. We propose a feature selection algorithm called Sequential Attention that achieves state-of-the-art empirical results for neural networks. This algorithm is based on an efficient one-pass implementation of greedy forward selection and uses attention weights at each step as a proxy for feature importance. We give theoretical insights into our algorithm for linear regression by showing that an adaptation to this setting is equivalent to the classical Orthogonal Matching Pursuit (OMP) algorithm, and thus inherits all of its provable guarantees. Our theoretical and empirical analyses offer new explanations towards the effectiveness of attention and its connections to overparameterization, which may be of independent interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f83ac23f-43da-4b0d-98fa-4727c49c7019Cited by top-tier papers5
- Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention LandscapeJuno Kim, Taiji SuzukiICML 2024 · 42 citations
- Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML SystemsBenjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang et al.NeurIPS 2023 · 30 citations
- Understanding MLP-Mixer as a wide and sparse MLPTomohiro Hayase, Ryo KarakidaICML 2024 · 9 citations
- SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial OptimizationTaisuke Yasuda, Kyriakos Axiotis, Gang Fu, Mohammad Hossein Bateni et al.NeurIPS 2024 · 1 citation
- SAND: One-Shot Feature Selection with Additive Noise DistortionPedram Pad, Hadi Hammoud, Mohamad Dia, Nadim Maamari et al.ICML 2025
Builds on6
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- Feature Importance Ranking for Deep LearningMaksymilian Wojtas, Ke ChenNeurIPS 2020 · 159 citations
- Feature Selection using Stochastic GatesYutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval KlugerICML 2020 · 39 citations
- Data-Efficient Structured Pruning via Submodular OptimizationMarwa El Halabi, Suraj Srinivas, Simon Lacoste-JulienNeurIPS 2022 · 31 citations
Related papers
- Sparse Bayesian Learning via Stepwise RegressionSebastian E. Ament, Carla P. GomesICML 2021 · 11 citations
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams et al.NeurIPS 2025 · 54 citations
- Information Maximization Perspective of Orthogonal Matching Pursuit with Applications to Explainable AIAditya Chattopadhyay, Ryan Pilgrim, René VidalNeurIPS 2023 · 17 citations
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De et al.ICML 2021 · 305 citations
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 18 citations
