Joint Model and Data Sparsification via the Marginal Likelihood
Alexander Timans, Thomas Moellenhoff, Christian Andersson Naesseth, Mohammad Emtiyaz Khan, Eric Nalisnick
Abstract
Sparse recovery in linear systems underpins applications from signal processing to high-dimensional regression. Sparse Bayesian Learning, grounded in the principle of automatic relevance determination (ARD), offers a practical Bayesian mechanism for feature sparsity via marginal likelihood optimization. Yet, its reliance on a homoscedastic noise model renders it sensitive to data contaminations such as outliers or misspecified noise, harming model fit and predictions. Instead, we propose jointly learning individual feature and sample relevancies, enabling simultaneous model and data sparsification via a single Bayesian objective. This symmetric pruning of model and data offers a natural extension that preserves conjugacy, admits closed-form updates for standard optimization procedures, and aligns with perspectives from robust regression and influence functions. Empirical results across diverse regression tasks affirm that a joint ARD approach consistently yields both sparse and robust prediction models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1105d89d-8d7e-44a4-97a2-de3fbe1a6d2dBuilds on12
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum et al.ICML 2022 · 83 citations
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- Effective Bayesian Heteroscedastic Regression with Deep Neural NetworksAlexander Immer, Emanuele Palumbo, Alexander Marx, Julia E. VogtNeurIPS 2023 · 34 citations
- Transductive Active Learning: Theory and ApplicationsJonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As et al.NeurIPS 2024 · 24 citations
Related papers
- Robust Gaussian Processes via Relevance PursuitSebastian Ament, Elizabeth Santorella, David Eriksson, Ben Letham et al.NeurIPS 2024 · 12 citations
- Efficient Network Automatic Relevance DeterminationHongwei Zhang, Ziqi Ye, Xinyuan Wang, Xin Guo et al.ICML 2025
- Sparse Bayesian Learning via Stepwise RegressionSebastian E. Ament, Carla P. GomesICML 2021 · 11 citations
- Robust and Computation-Aware Gaussian ProcessesMarshal Arijona Sinaga, Julien Martinelli, Samuel KaskiNeurIPS 2025 · 1 citation
- Robust and Conjugate Gaussian Process RegressionMatías Altamirano, François-Xavier Briol, Jeremias KnoblauchICML 2024 · 18 citations
