Efficient Parametric Approximations of Neural Network Function Space Distance
Nikita Dhawan, Sicong Huang, Juhan Bae, Roger Baker Grosse
Abstract
It is often useful to compactly summarize important properties of model parameters and training data so that they can be used later without storing and/or iterating over the entire dataset. As a specific case, we consider estimating the Function Space Distance (FSD) over a training set, i.e. the average discrepancy between the outputs of two neural networks. We propose a Linearized Activation Function TRick (LAFTR) and derive an efficient approximation to FSD for ReLU neural networks. The key idea is to approximate the architecture as a linear network with stochastic gating. Despite requiring only one parameter per unit of the network, our approach outcompetes other parametric approximations with larger memory requirements. Applied to continual learning, our parametric approximation is competitive with state-of-the-art nonparametric approximations, which require storing many training examples. Furthermore, we show its efficacy in estimating influence functions accurately and detecting mislabeled examples without expensive iterations over the entire dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea465f60-efd7-4952-93cc-dd4d69c69859Cited by top-tier papers5
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng et al.NeurIPS 2024 · 70 citations
- On the Diminishing Returns of Width for Continual LearningEtash Kumar Guha, Vihan LakshmanICML 2024 · 9 citations
- Dataless Weight Disentanglement in Task Arithmetic via Kronecker-Factored Approximate CurvatureAngelo Porrello, Pietro Buzzega, Felix Dangel, Thomas Sommariva et al.ICLR 2026 · 6 citations
- Preserving Linear Separability in Continual Learning by Backward Feature ProjectionQiao Gu, Dongsub Shim, Florian ShkurtiCVPR 2023
- Streamlining Prediction in Bayesian Deep LearningRui Li, Marcus Klasson, Arno Solin, Martin TrappICLR 2025
Builds on17
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 211 citations
- Functional Regularisation for Continual Learning with Gaussian ProcessesMichalis K. Titsias, Jonathan Schwarz, Alexander G. de G. Matthews, Razvan Pascanu et al.ICLR 2020 · 209 citations
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
Related papers
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
- Neural approximation of Wasserstein distance via a universal architecture for symmetric and factorwise group invariant functionsSamantha Chen, Yusu WangNeurIPS 2023 · 4 citations
- Estimating informativeness of samples with Smooth Unique InformationHrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder et al.ICLR 2021 · 26 citations
- Inner Product-based Neural Network SimilarityWei Chen, Zichen Miao, Qiang QiuNeurIPS 2023 · 14 citations
- Exploiting Space Folding by Neural NetworksMichal Lewandowski, Raphael Pisoni, Bernhard Heinzl, Bernhard Alois MoserAAAI 2026
