ModHiFi: Identifying High Fidelity predictive components for Model Modification
Dhruva Kashyap, Chaitanya Murti, Pranav K. Nayak, Tanay Narshana, Chiranjib Bhattacharyya
Abstract
Open weight models, which are ubiquitous, rarely provide access to their training data or loss function. This makes modifying such models for tasks such as pruning or unlearning, which are constrained by this unavailability, an active area of research. Existing techniques typically require gradients or ground-truth labels, rendering them infeasible in settings with limited computational resources. In this work, we investigate the fundamental question of identifying components that are critical to the model's predictive performance, without access to either gradients or the loss function, and with only distributional access such as synthetic data. We theoretically demonstrate that the global error is linearly bounded by local reconstruction errors for Lipschitz-continuous networks such as CNNs and well-trained Transformers (which, contrary to existing literature, we find exhibit Lipschitz continuity). This motivates using the locally reconstructive behavior of component subsets to quantify their global importance, via a metric that we term Subset Fidelity. In the uncorrelated features setting, selecting individual components based on their Subset Fidelity scores is optimal, which we utilize to propose ModHiFi, an algorithm for model modification that requires neither training data nor access to a loss function. ModHiFi-P, for structured pruning, achieves an 11% speedup over the current state of the art on ImageNet models and competitive performance on language models. ModHiFi-U, for classwise unlearning, achieves complete unlearning on CIFAR-10 without fine-tuning and demonstrates competitive performance on Swin Transformers. 2 * Author primarily contributed to this work before joining Google. 2 Our code is available at https://github.com/DhruvaKashyap/modhifi 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
related to privacy and security [88], and also the use of synthetic data, which has become critical in a variety of language modeling settings [10,72]. Thus, we address the challenging problem of altering well-trained models without training data or the loss function, and only with distributional access to the original training distribution in the form of synthetic data, focusing specifically on structured pruning and classwise unlearning.
Modifying open weight models without the loss function and only synthetic data requires answering a fundamental question: which components in a model contribute significantly to its predictive performance 3 ? However, most methods that identify critical components for specific modifications (e.g., pruning) cannot be applied to others (e.g., unlearning) [53], often require expensive fine-tuning, and are architecture-specific. Moreover, most methods utilize gradients to assess the impact of a component on the loss objective, which is not feasible in the absence of the loss function and the training data. While the LLM pruning literature uses calibration datasets to mitigate the problem of the absence of datasets [2,46], the problem of achieving sparsity in vision models without original training data is hard and unsolved [27]. Moreover, the issue of performing classwise unlearning without access to the original training data has not been addressed [32,53].
Towards enabling the modification of well-trained open weight models amidst these challenges, we make the following contributions:
(C1) Local-to-Global with Lipschitzness. An open question is the extent to which local model modifications impact the predictive performance of the model. In the absence of loss functions and training sets, estimating the impact of component modification by using gradients (as done in [32,43,46]) is infeasible. To address this, in Theorem 3.6, we show that for Lipschitz continuous networks, the reconstruction error at the final layer is at most linear in the local reconstruction errors. Moreover, contrary to the assertion that transformers are not Lipschitz continuous [60], in Corollary B.4, we show that this is not the case for well-trained transformers, allowing us to apply Theorem 3.6 to not just CNNs, but well-trained ViTs and LLMs as well.
(C2) Identifying Subsets of Important Components. Contrary to prior work, which usually infers saliencies for single components, we propose measuring the importance of sets of components to understand the cumulative effects of groups of components on a model's predictive performance.
Leveraging Theorem 3.6, we propose Subset Fidelity, which quantifies the extent to which a subset of components can reconstruct the output after modifying their weights. However, computing optimal subsets is NP-complete, motivating us to compute Subset Fidelity scores for singleton sets. Theorem 3.9 establishes that selecting singletons with the highest subset fidelity scores is optimal when the features are uncorrelated.
(C3) Modifying Models with ModHiFi-X. Motivated by Theorem 3.9, we propose the ModHiFi algorithm, which uses the subset
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ed889a4-134a-43c9-a5f8-7aa5fcfcb54aBuilds on45
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
Related papers
- DisCEdit: Model Editing by Identifying Discriminative ComponentsChaitanya Murti, Chiranjib BhattacharyyaNeurIPS 2024 · 2 citations
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningRuchika Chavhan, Da Li, Timothy M. HospedalesICLR 2025 · 2 citations
- Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion ModelsCi Zhang, Zhaojun Ding, Chence Yang, Jun Liu et al.CVPR 2026 · 1 citation
- Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language ModelsZheyuan Liu, Guangyao Dou, Xiangchi Yuan, Chunhui Zhang et al.ACL 2025
- Stealix: Model Stealing via Prompt EvolutionZhixiong Zhuang, Hui-Po Wang, Maria-Irina Nicolae, Mario FritzICML 2025
