RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models
Pingyi Hu, Xiaofan Bai, Xiaojing Ma, Chaoxiang He, Dongmei Zhang, Bin Benjamin Zhu
摘要
The proliferation of Machine Learning as a Service (MLaaS) has enabled widespread deployment of large language models (LLMs) via cloud APIs, but also raises critical concerns about model integrity and security. Existing black-box tamper detection methods, such as watermarking and fingerprinting, rely on the stability of model outputs-a property that does not hold for inherently stochastic LLMs. We address this challenge by formulating blackbox tamper detection for LLMs as a hypothesistesting problem. To enable efficient and sensitive fingerprinting, we derive a first-order surrogate for KL divergence-the entropy-gradient norm-to identify prompts most responsive to parameter perturbations. Building on this, we propose Regularized Entropy-Sensitive Fingerprinting (RESF), which enhances sensitivity while regularizing entropy to improve output stability and control false positives. To further distinguish tampering from benign randomness, such as temperature shifts, RESF employs a lightweight two-tier sequential test combining support-based and distributional checks with rigorous false-alarm control. Comprehensive analysis and experiments across multiple LLMs show that RESF achieves up to 98.80% detection accuracy under challenging conditions, such as minimal LoRA fine-tuning with five optimized fingerprints. RESF consistently demonstrates strong sensitivity and robustness, providing an effective and scalable solution for black-box tamper detection in cloud-deployed LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language ModelsRuisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, Farinaz KoushanfarUSENIX Security 2024 · 被引用 88 次
- Protecting Intellectual Property of Large Language Model-Based Code Generation APIs via WatermarksZongjie Li, Chaozheng Wang, Shuai Wang, Cuiyun GaoCCS 2023 · 被引用 25 次
- MetaV: A Meta-Verifier Approach to Task-Agnostic Model FingerprintingXudong Pan, Yifan Yan, Mi Zhang, Min YangKDD 2022 · 被引用 19 次
- AID: Attesting the Integrity of Deep Neural NetworksOmid Aramoon, Pin-Yu Chen, Gang QuDAC 2021 · 被引用 9 次
- Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN ModelsXiaofan Bai, Chaoxiang He, Xiaojing Ma, Bin Benjamin Zhu 等ICML 2024 · 被引用 6 次
相关 Paper
- Towards Stricter Black-box Integrity Verification of Deep Neural Network ModelsChaoxiang He, Xiaofan Bai, Xiaojing Ma, Bin B. Zhu 等ACM MM 2024 · 被引用 3 次
- LLM Fingerprinting via Semantically Conditioned WatermarksThibaud Gloaguen, Robin Staab, Nikola Jovanovic, Martin T. VechevICLR 2026 · 被引用 8 次
- PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning DetectionJungmin Lee, Peizhuo Lv, Yeonjoon LeeACL 2026
- ImF: Embedding an Implicit Fingerprint in Your Large Language ModelsJiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue 等ACL 2026
- Watermarking Large Language Models: An Unbiased and Low-risk MethodMinjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang 等ACL 2025 · 被引用 6 次
