LLM Fingerprinting via Semantically Conditioned Watermarks
Thibaud Gloaguen, Robin Staab, Nikola Jovanovic, Martin T. Vechev
Abstract
Most LLM fingerprinting methods teach the model to respond to a few fixed queries with predefined atypical responses (keys). This memorization often does not survive common deployment steps such as finetuning or quantization, and such keys can be easily detected and filtered from LLM responses, ultimately breaking the fingerprint. To overcome these limitations we introduce LLM fingerprinting via semantically conditioned watermarks, replacing fixed query sets with a broad semantic domain, and replacing brittle atypical keys with a statistical watermarking signal diffused throughout each response. After teaching the model to watermark its responses only to prompts from a predetermined domain e.g., French language, the model owner can use queries from that domain to reliably detect the fingerprint and verify ownership. As we confirm in our thorough experimental evaluation, our fingerprint is both stealthy and robust to all common deployment scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb4106e9-4d4e-4ee7-902c-e8430e7a0ffeCited by top-tier papers1
Ask how each one uses itBuilds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
Related papers
- SIF: Semantically In-Distribution Fingerprints for Large Vision-Language ModelsYifei Zhao, Qian Lou, Mengxin ZhengCVPR 2026 · 2 citations
- ImF: Embedding an Implicit Fingerprint in Your Large Language ModelsJiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue et al.ACL 2026
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse AutoencodersZhuohao Yu, Xingru Jiang, Weizheng Gu, Yidong Wang et al.NeurIPS 2025 · 6 citations
- UMMF: Protecting Copyright of Large Vision-Language Models through Unlearning-based Multimodal Memorization FingerprintXiaofan Zheng, Xinghao Wang, Xiaojun WanACL 2026
- Segmenting Watermarked Texts From Language ModelsXingchi Li, Guanxun Li, Xianyang ZhangNeurIPS 2024 · 5 citations
