MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic Perspective
Xinwei Zhang, Haibo Hu, Qingqing Ye, Li Bai, Huadi Zheng
Abstract
Information leakage issues in machine learning-based Web applications have attracted increasing attention. While the risk of data privacy leakage has been rigorously analyzed, the theory of model function leakage, known as Model Extraction Attacks (MEAs), has not been well studied. In this paper, we are the first to understand MEAs theoretically from an attack-agnostic perspective and to propose analytical metrics for evaluating model extraction risks. By using the Neural Tangent Kernel (NTK) theory, we formulate the linearized MEA as a regularized kernel classification problem and then derive the fidelity gap and generalization error bounds of the attack performance. Based on these theoretical analyses, we propose a new theoretical metric called Model Recovery Complexity (MRC), which measures the distance of weight changes between the victim and surrogate models to quantify risk. Additionally, we find that victim model accuracy, which shows a strong positive correlation with model extraction risk, can serve as an empirical metric. By integrating these two metrics, we propose a framework, namely Model Extraction Risk Inspector (MER-Inspector), to compare the extraction risks of models under different model architectures by utilizing relative metric values. We conduct extensive experiments on 16 model architectures and 5 datasets. The experimental results demonstrate that the proposed metrics have a high correlation with model extraction risks, and MER-Inspector can accurately compare the extraction risks of any two models with up to 89.58%. CCS Concepts • Security and privacy → Web application security; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1855fb2f-19a5-4a9f-9665-682fb5569ba5Cited by top-tier papers5
- On the Adversarial Robustness of Large Vision-Language Models under Visual Token CompressionXinwei Zhang, Hangcheng Liu, Li Bai, Hao Wang et al.ICML 2026 · 2 citations
- DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor AttacksJiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo et al.AAAI 2026
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction AttacksYaxin Xiao, Qingqing Ye, Zi Liang, Haoyang Li et al.AAAI 2026
- SLAT: Segment-Level Adaptive Trimming for Efficient CoT ReasoningJian Yao, Xiongcai Luo, Ran Cheng, KC TanICML 2026
- Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao et al.ICML 2025
Builds on27
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock et al.ICCV 2021 · 1,009 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Systematic Evaluation of Privacy Risks of Machine Learning ModelsLiwei Song, Prateek MittalUSENIX Security 2021 · 483 citations
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade et al.AAAI 2020 · 164 citations
Related papers
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 3 citations
- High Accuracy and High Fidelity Extraction of Neural NetworksMatthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin et al.USENIX Security 2020
- CREDIT: Certified Ownership Verification of Deep Neural Networks Against Model Extraction AttacksBolin Shen, Zhan Cheng, Neil Gong, Fan Yao et al.ICML 2026 · 3 citations
- ML-Doctor: Holistic Risk Assessment of Inference Attacks Against Machine Learning ModelsYugeng Liu, Rui Wen, Xinlei He, Ahmed Salem et al.USENIX Security 2022
- Towards Model Extraction Attacks in GAN-Based Image Translation via Domain Shift MitigationDi Mi, Yanjun Zhang, Leo Yu Zhang, Shengshan Hu et al.AAAI 2024 · 5 citations
