Hyperparameter Optimization through Neural Network Partitioning
Bruno Mlodozeniec, Matthias Reisser, Christos Louizos
摘要
Well-tuned hyperparameters are crucial for obtaining good generalization behavior in neural networks. They can enforce appropriate inductive biases, regularize the model and improve performance -- especially in the presence of limited data. In this work, we propose a simple and efficient way for optimizing hyperparameters inspired by the marginal likelihood, an optimization objective that requires no validation data. Our method partitions the training data and a neural network model into data shards and parameter partitions, respectively. Each partition is associated with and optimized only on specific data shards. Combining these partitions into subnetworks allows us to define the ``out-of-training-sample" loss of a subnetwork, i.e., the loss on data shards unseen by the subnetwork, as the objective for hyperparameter optimization. We demonstrate that we can apply this objective to optimize a variety of different hyperparameters in a single training run while being significantly computationally cheaper than alternative methods aiming to optimize the marginal likelihood for neural networks. Lastly, we also focus on optimizing hyperparameters in federated learning, where retraining and cross-validation are particularly challenging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning Layer-wise Equivariances Automatically using GradientsTycho F. A. van der Ouderaa, Alexander Immer, Mark van der WilkNeurIPS 2023 · 被引用 28 次
- Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsAlexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch 等ICML 2023 · 被引用 17 次
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán 等NeurIPS 2024 · 被引用 16 次
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger 等NeurIPS 2025
- Diffusion-based Neural Network Weights GenerationBedionita Soro, Bruno Andreis, Hayeon Lee, Wonyong Jeong 等ICLR 2025
它引用的顶会 Paper6
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li 等NeurIPS 2021 · 被引用 111 次
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum 等ICML 2022 · 被引用 83 次
- Learning Invariances in Neural Networks from Training DataGregory W. Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 被引用 78 次
相关 Paper
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang 等VLDB 2022 · 被引用 42 次
- FedPop: Federated Population-based Hyperparameter TuningHaokun Chen, Denis Krompaß, Jindong Gu, Volker TrespAAAI 2025 · 被引用 3 次
- Single-shot General Hyper-parameter Optimization for Federated LearningYi Zhou, Parikshit Ram, Theodoros Salonidis, Nathalie Baracaldo 等ICLR 2023 · 被引用 1 次
- Vertical Federated Learning with Missing Features During Training and InferencePedro Valdeira, Shiqiang Wang, Yuejie ChiICLR 2025
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 被引用 452 次
