RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models
Pingyi Hu, Xiaofan Bai, Xiaojing Ma, Chaoxiang He, Dongmei Zhang, Bin Benjamin Zhu
Abstract
The proliferation of Machine Learning as a Service (MLaaS) has enabled widespread deployment of large language models (LLMs) via cloud APIs, but also raises critical concerns about model integrity and security. Existing black-box tamper detection methods, such as watermarking and fingerprinting, rely on the stability of model outputs-a property that does not hold for inherently stochastic LLMs. We address this challenge by formulating blackbox tamper detection for LLMs as a hypothesistesting problem. To enable efficient and sensitive fingerprinting, we derive a first-order surrogate for KL divergence-the entropy-gradient norm-to identify prompts most responsive to parameter perturbations. Building on this, we propose Regularized Entropy-Sensitive Fingerprinting (RESF), which enhances sensitivity while regularizing entropy to improve output stability and control false positives. To further distinguish tampering from benign randomness, such as temperature shifts, RESF employs a lightweight two-tier sequential test combining support-based and distributional checks with rigorous false-alarm control. Comprehensive analysis and experiments across multiple LLMs show that RESF achieves up to 98.80% detection accuracy under challenging conditions, such as minimal LoRA fine-tuning with five optimized fingerprints. RESF consistently demonstrates strong sensitivity and robustness, providing an effective and scalable solution for black-box tamper detection in cloud-deployed LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language ModelsRuisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, Farinaz KoushanfarUSENIX Security 2024 · 88 citations
- Protecting Intellectual Property of Large Language Model-Based Code Generation APIs via WatermarksZongjie Li, Chaozheng Wang, Shuai Wang, Cuiyun GaoCCS 2023 · 25 citations
- MetaV: A Meta-Verifier Approach to Task-Agnostic Model FingerprintingXudong Pan, Yifan Yan, Mi Zhang, Min YangKDD 2022 · 19 citations
- AID: Attesting the Integrity of Deep Neural NetworksOmid Aramoon, Pin-Yu Chen, Gang QuDAC 2021 · 9 citations
- Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN ModelsXiaofan Bai, Chaoxiang He, Xiaojing Ma, Bin Benjamin Zhu et al.ICML 2024 · 6 citations
Related papers
- Towards Stricter Black-box Integrity Verification of Deep Neural Network ModelsChaoxiang He, Xiaofan Bai, Xiaojing Ma, Bin B. Zhu et al.ACM MM 2024 · 3 citations
- LLM Fingerprinting via Semantically Conditioned WatermarksThibaud Gloaguen, Robin Staab, Nikola Jovanovic, Martin T. VechevICLR 2026 · 8 citations
- PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning DetectionJungmin Lee, Peizhuo Lv, Yeonjoon LeeACL 2026
- ImF: Embedding an Implicit Fingerprint in Your Large Language ModelsJiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue et al.ACL 2026
- Watermarking Large Language Models: An Unbiased and Low-risk MethodMinjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang et al.ACL 2025 · 6 citations
