Sequential Attention for Feature Selection
Taisuke Yasuda, Mohammad Hossein Bateni, Lin Chen, Matthew Fahrbach, Gang Fu, Vahab Mirrokni
摘要
Feature selection is the problem of selecting a subset of features for a machine learning model that maximizes model quality subject to a budget constraint. For neural networks, prior methods, including those based on regularization, attention, and other techniques, typically select the entire feature subset in one evaluation round, ignoring the residual value of features during selection, i.e., the marginal contribution of a feature given that other features have already been selected. We propose a feature selection algorithm called Sequential Attention that achieves state-of-the-art empirical results for neural networks. This algorithm is based on an efficient one-pass implementation of greedy forward selection and uses attention weights at each step as a proxy for feature importance. We give theoretical insights into our algorithm for linear regression by showing that an adaptation to this setting is equivalent to the classical Orthogonal Matching Pursuit (OMP) algorithm, and thus inherits all of its provable guarantees. Our theoretical and empirical analyses offer new explanations towards the effectiveness of attention and its connections to overparameterization, which may be of independent interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention LandscapeJuno Kim, Taiji SuzukiICML 2024 · 被引用 42 次
- Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML SystemsBenjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang 等NeurIPS 2023 · 被引用 30 次
- Understanding MLP-Mixer as a wide and sparse MLPTomohiro Hayase, Ryo KarakidaICML 2024 · 被引用 9 次
- SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial OptimizationTaisuke Yasuda, Kyriakos Axiotis, Gang Fu, Mohammad Hossein Bateni 等NeurIPS 2024 · 被引用 1 次
- SAND: One-Shot Feature Selection with Additive Noise DistortionPedram Pad, Hadi Hammoud, Mohamad Dia, Nadim Maamari 等ICML 2025
它引用的顶会 Paper6
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Feature Importance Ranking for Deep LearningMaksymilian Wojtas, Ke ChenNeurIPS 2020 · 被引用 159 次
- Feature Selection using Stochastic GatesYutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval KlugerICML 2020 · 被引用 39 次
- Data-Efficient Structured Pruning via Submodular OptimizationMarwa El Halabi, Suraj Srinivas, Simon Lacoste-JulienNeurIPS 2022 · 被引用 31 次
相关 Paper
- Sparse Bayesian Learning via Stepwise RegressionSebastian E. Ament, Carla P. GomesICML 2021 · 被引用 11 次
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 等NeurIPS 2025 · 被引用 54 次
- Information Maximization Perspective of Orthogonal Matching Pursuit with Applications to Explainable AIAditya Chattopadhyay, Ryan Pilgrim, René VidalNeurIPS 2023 · 被引用 17 次
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De 等ICML 2021 · 被引用 305 次
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 被引用 18 次
