TPV: Parameter Perturbations Through the Lens of Test Prediction Variance
Devansh Arpit
Abstract
We introduce test prediction variance (TPV)—the first-order sensitivity of a trained model's outputs to parameter perturbations—as a unifying framework for analyzing post-training robustness. TPV's trace form separates the geometry of the trained model from the perturbation covariance , placing SGD noise, label noise, quantization, and pruning under a single lens. The resulting expressions recover the wide-minima hypothesis for SGD and quantization noise, and yield a distinct Jacobian-spectral characterization for label noise connecting label-noise TPV with benign overfitting in nonlinear networks. Theoretically, we prove that training-set TPV converges to its test-set counterpart in the overparameterized limit, irrespective of generalization performance, providing the first result that prediction variance under local parameter perturbations can be inferred from training inputs alone. Empirically, this stability holds far more broadly, including at very low widths. Further, TPV correlates well with test loss, enabling practical applications: JBR, a label-free pruning criterion derived from TPV geometry matching state-of-the-art baselines; and training-set based model selection signal for in-distribution and transfer learning scenarios. Code Available Here
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd4207a2-a04d-4342-9788-0b5e69aeec22Builds on9
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Group Fisher Pruning for Practical Network CompressionLiyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou et al.ICML 2021 · 204 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
Related papers
- Training-Free Uncertainty Estimation for Dense Regression: Sensitivity as a SurrogateLu Mi, Hao Wang, Yonglong Tian, Hao He et al.AAAI 2022 · 36 citations
- The Memory-Perturbation Equation: Understanding Model's Sensitivity to DataPeter Nickl, Lu Xu, Dharmesh Tailor, Thomas Möllenhoff et al.NeurIPS 2023 · 17 citations
- The Generalization-Stability Tradeoff In Neural Network PruningBrian R. Bartoldson, Ari S. Morcos, Adrian Barbu, Gordon ErlebacherNeurIPS 2020 · 97 citations
- Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter CorruptionXu Sun, Zhiyuan Zhang, Xuancheng Ren, Ruixuan Luo et al.AAAI 2021 · 45 citations
- How Benign is Benign Overfitting ?Amartya Sanyal, Puneet K. Dokania, Varun Kanade, Philip H. S. TorrICLR 2021 · 61 citations
