A Zest of LIME: Towards Architecture-Independent Model Distances
Hengrui Jia, Hongyu Chen, Jonas Guan, Ali Shahin Shamsabadi, Nicolas Papernot
Abstract
Definitions of the distance between two machine learning models either characterize the similarity of the models' predictions or of their weights. While similarity of weights is attractive because it implies similarity of predictions in the limit, it suffers from being inapplicable to comparing models with different architectures. On the other hand, the similarity of predictions is broadly applicable but depends heavily on the choice of model inputs during comparison. In this paper, we instead propose to compute distance between black-box models by comparing their Local Interpretable Model-Agnostic Explanations (LIME). To compare two models, we take a reference dataset, and locally approximate the models on each reference point with linear models trained by LIME. We then compute the cosine distance between the concatenated weights of the linear models. This yields an approach that is both architecture-independent and possesses the benefits of comparing models in weight space. We empirically show that our method, which we call Zest, can be applied to two problems that require measurements of model similarity: detecting model stealing and machine unlearning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 75d66c4a-d7e4-4117-aa36-b7ca382d8288Cited by top-tier papers11
- DAWN: Dynamic Adversarial Watermarking of Neural NetworksSebastian Szyller, Buse Gul Atli, Samuel Marchal, N. AsokanACM MM 2021 · 133 citations
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan et al.NSDI 2023 · 94 citations
- Model Provenance Testing for Large Language ModelsIvica Nikolic, Teodora Baluta, Prateek SaxenaNeurIPS 2025 · 20 citations
- United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesTianlong Xu, Chen Wang, Gaoyang Liu, Yang Yang et al.NeurIPS 2024 · 17 citations
- ModelGiF: Gradient Fields for Model Functional DistanceJie Song, Zhengqi Xu, Sai Wu, Gang Chen et al.ICCV 2023 · 6 citations
Related papers
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 56 citations
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 7 citations
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
- Sparse and Faithful Local Explanations with Piecewise Linear SurrogatesYixin Wang, Yucheng DongICML 2026
- Learning to Explain: Generating Stable Explanations FastXuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf et al.ACL 2021
