Indirectly Parameterized Concrete Autoencoders
Alfred Nilsson, Klas Wijk, Sai Bharath Chandra Gutha, Erik Englesson, Alexandra Hotti, Carlo Saccardi, Oskar Kviman, Jens Lagergren, Ricardo Vinuesa, Hossein Azizpour
Abstract
Feature selection is a crucial task in settings where data is high-dimensional or acquiring the full set of features is costly. Recent developments in neural network-based embedded feature selection show promising results across a wide range of applications. Concrete Autoencoders (CAEs), considered state-of-the-art in embedded feature selection, may struggle to achieve stable joint optimization, hurting their training time and generalization. In this work, we identify that this instability is correlated with the CAE learning duplicate selections. To remedy this, we propose a simple and effective improvement: Indirectly Parameterized CAEs (IP-CAEs). IP-CAEs learn an embedding and a mapping from it to the Gumbel-Softmax distributions' parameters. Despite being simple to implement, IP-CAE exhibits significant and consistent improvements over CAE in both generalization and training time across several datasets for reconstruction and classification. Unlike CAE, IP-CAE effectively leverages non-linear relationships and does not require retraining the jointly optimized decoder. Furthermore, our approach is, in principle, generalizable to Gumbel-Softmax distributions beyond feature selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec2f7123-2d0a-495f-b953-317670cd9fffCited by top-tier papers1
Ask how each one uses itBuilds on8
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 170 citations
- Experimental design for MRI by greedy policy searchTim Bakker, Herke van Hoof, Max WellingNeurIPS 2020 · 70 citations
- Deep probabilistic subsampling for task-adaptive compressed sensingIris A. M. Huijben, Bastiaan S. Veeling, Ruud J. G. van SlounICLR 2020 · 47 citations
- Feature Selection using Stochastic GatesYutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval KlugerICML 2020 · 39 citations
Related papers
- Discrete Variational Autoencoding via Policy SearchMichael Drolet, Firas Al-Hafez, Aditya Bhatt, Jan Peters et al.ICLR 2026
- Latent Template Induction with Gumbel-CRFsYao Fu, Chuanqi Tan, Bin Bi, Mosha Chen et al.NeurIPS 2020 · 15 citations
- Toward Identifiable Sparse AutoencodersWalter Nelson, Theofanis Karaletsos, Francesco LocatelloICML 2026 · 1 citation
- Ensembling Sparse AutoencodersSoham Gadgil, Chris Lin, Su-In LeeICML 2026
- A Unified Deep Model of Learning from both Data and Queries for Cardinality EstimationPeizhi Wu, Gao CongSIGMOD 2021 · 73 citations
