FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences
Richardeau Gurvan, Gohar Dashyan, Erwan Le Merrer, Gilles Tredan
摘要
Literature reveals that a Large Language Model's (LLM) behavior is not only conditioned by its original weights but also its instance-level parameters, such as instructional prompt, sampling configuration or quantization. A model that generates safe outputs under one configuration may produce toxic content under another. However, current LLM identification techniques (such as fingerprinting) focus on intellectual property protection, and their design favors robustness to changes in these instance-level parameters. This poses a critical challenge for AI regulation in which compliance assessments target actual deployed behaviors, not model provenance. In this paper, we introduce instance-level fingerprinting, a regulator-oriented paradigm that distinguishes configurations of the same LLM. Our method FLIPS, exploits biases in generated binary random sequences to reach 96% (closed-set) and 90% (open-set, where some targets are unknown) identification accuracy across 237 model instances, versus 35% for the adapted LLMmap baseline. This shows that instance-level fingerprinting is both necessary for regulation and practically feasible. Code available at https://github.com/GurvanR/FLIPS-LLM-Instance-Fingerprinting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger 等ICLR 2024 · 被引用 373 次
- Fingerprinting Deep Neural Networks Globally via Universal Adversarial PerturbationsZirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang 等CVPR 2022 · 被引用 66 次
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
相关 Paper
- LLMmap: Fingerprinting for Large Language ModelsDario Pasquini, Evgenios M. Kornaropoulos, Giuseppe AtenieseUSENIX Security 2025
- PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning DetectionJungmin Lee, Peizhuo Lv, Yeonjoon LeeACL 2026
- Quantifying Large Language Model Attacks Through the Lens of Model CognitionXiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai 等USENIX Security 2026
- ImF: Embedding an Implicit Fingerprint in Your Large Language ModelsJiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue 等ACL 2026
- Fingerprinting LLMs via Prompt InjectionYuepeng Hu, Zhengyuan Jiang, Mengyuan Li, Osama Ahmed 等ACL 2026 · 被引用 3 次
