Hyperparameter Optimization through Neural Network Partitioning
Bruno Mlodozeniec, Matthias Reisser, Christos Louizos
Abstract
Well-tuned hyperparameters are crucial for obtaining good generalization behavior in neural networks. They can enforce appropriate inductive biases, regularize the model and improve performance -- especially in the presence of limited data. In this work, we propose a simple and efficient way for optimizing hyperparameters inspired by the marginal likelihood, an optimization objective that requires no validation data. Our method partitions the training data and a neural network model into data shards and parameter partitions, respectively. Each partition is associated with and optimized only on specific data shards. Combining these partitions into subnetworks allows us to define the ``out-of-training-sample" loss of a subnetwork, i.e., the loss on data shards unseen by the subnetwork, as the objective for hyperparameter optimization. We demonstrate that we can apply this objective to optimize a variety of different hyperparameters in a single training run while being significantly computationally cheaper than alternative methods aiming to optimize the marginal likelihood for neural networks. Lastly, we also focus on optimizing hyperparameters in federated learning, where retraining and cross-validation are particularly challenging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ec8cc6d-7c54-44ba-8f68-8f0a2bfffa2eCited by top-tier papers5
- Learning Layer-wise Equivariances Automatically using GradientsTycho F. A. van der Ouderaa, Alexander Immer, Mark van der WilkNeurIPS 2023 · 28 citations
- Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsAlexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch et al.ICML 2023 · 17 citations
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán et al.NeurIPS 2024 · 16 citations
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger et al.NeurIPS 2025
- Diffusion-based Neural Network Weights GenerationBedionita Soro, Bruno Andreis, Hayeon Lee, Wonyong Jeong et al.ICLR 2025
Builds on6
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li et al.NeurIPS 2021 · 111 citations
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum et al.ICML 2022 · 83 citations
- Learning Invariances in Neural Networks from Training DataGregory W. Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 78 citations
Related papers
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang et al.VLDB 2022 · 42 citations
- FedPop: Federated Population-based Hyperparameter TuningHaokun Chen, Denis Krompaß, Jindong Gu, Volker TrespAAAI 2025 · 3 citations
- Single-shot General Hyper-parameter Optimization for Federated LearningYi Zhou, Parikshit Ram, Theodoros Salonidis, Nathalie Baracaldo et al.ICLR 2023 · 1 citation
- Vertical Federated Learning with Missing Features During Training and InferencePedro Valdeira, Shiqiang Wang, Yuejie ChiICLR 2025
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 452 citations
