PEP: Parameter Ensembling by Perturbation
Alireza Mehrtash, Purang Abolmaesumi, Polina Golland, Tina Kapur, Demian Wassermann, William M. Wells III
摘要
Ensembling is now recognized as an effective approach for increasing the predictive performance and calibration of deep networks. We introduce a new approach, Parameter Ensembling by Perturbation (PEP), that constructs an ensemble of parameter values as random perturbations of the optimal parameter set from training by a Gaussian with a single variance parameter. The variance is chosen to maximize the log-likelihood of the ensemble average ( L ) on the validation data set. Empirically, and perhaps surprisingly, L has a well-defined maximum as the variance grows from zero (which corresponds to the baseline model). Conveniently, calibration level of predictions also tends to grow favorably until the peak of L is reached. In most experiments, PEP provides a small improvement in performance, and, in some cases, a substantial improvement in empirical calibration. We show that this "PEP effect" (the gain in log-likelihood) is related to the mean curvature of the likelihood function and the empirical Fisher information. Experiments on ImageNet pre-trained networks including ResNet, DenseNet, and Inception showed improved calibration and likelihood. We further observed a mild improvement in classification accuracy on these networks. Experiments on classification benchmarks such as MNIST and CIFAR-10 showed improved calibration and likelihood, as well as the relationship between the PEP effect and overfitting; this demonstrates that PEP can be used to probe the level of overfitting that occurred during training. In general, no special training procedure or network architecture is needed, and in the case of pre-trained networks, no additional training is needed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Triggering Failures: Out-Of-Distribution detection by learning from local adversarial attacks in Semantic SegmentationVictor Besnier, Andrei Bursuc, David Picard, Alexandre BriotICCV 2021 · 被引用 57 次
- Sparse Model Soups: A Recipe for Improved Pruning via Model AveragingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2024 · 被引用 22 次
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained WeightsYulu Gan, Phillip IsolaICML 2026 · 被引用 17 次
- Sharpness-diversity tradeoff: improving flat ensembles with SharpBalanceHaiquan Lu, Xiaotian Liu, Yefan Zhou, Qunli Li 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper2
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 被引用 344 次
- Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessAhmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger 等CVPR 2020
相关 Paper
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- EnsLoss: Stochastic Calibrated Loss Ensembles for Preventing Overfitting in ClassificationBen DaiICML 2025
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu 等ICLR 2021 · 被引用 235 次
- Post-Hoc Uncertainty Calibration for Domain Drift ScenariosChristian Tomani, Sebastian Gruber, Muhammed Ebrar Erdem, Daniel Cremers 等CVPR 2021
- Model Calibration in Dense Classification with Adaptive Label PerturbationJiawei Liu, Changkun Ye, Shan Wang, Ruikai Cui 等ICCV 2023 · 被引用 7 次
