Task-Agnostic Language Model Watermarking via High Entropy Passthrough Layers
Vaden Masrani, Mohammad Akbari, David Ming Xuan Yue, Ahmad Rezaei, Yong Zhang
摘要
In the era of costly pre-training of large language models, ensuring the intellectual property rights of model owners, and insuring that said models are responsibly deployed, is becoming increasingly important. To this end, we propose model watermarking via passthrough layers, which are added to existing pre-trained networks and trained using a self-supervised loss such that the model produces high-entropy output when prompted with a unique private key, and acts normally otherwise. Unlike existing model watermarking methods, our method is fully task-agnostic, and can be applied to both classification and sequence-to-sequence tasks without requiring advanced access to downstream fine-tuning datasets. We evaluate the proposed passthrough layers on a wide range of downstream tasks, and show experimentally our watermarking method achieves a near-perfect watermark extraction accuracy and false-positive rate in most cases without damaging original model performance. Additionally, we show our method is robust to both downstream fine-tuning, fine-pruning, and layer removal attacks, and can be trained in a fraction of the time required to train the original model. Code is available in the paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
- Simple linear attention language models balance the recall-throughput tradeoffSimran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina 等ICML 2024 · 被引用 154 次
- CATER: Intellectual Property Protection on Text Generation APIs via Conditional WatermarksXuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu 等NeurIPS 2022 · 被引用 106 次
相关 Paper
- PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant AttacksZhenxin Ai, Haiyun HeICML 2026 · 被引用 4 次
- Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model WatermarkingCong Kong, Rui Xu, Jiawei Chen, Zhaoxia YinACM MM 2025 · 被引用 1 次
- SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained EncodersTianshuo Cong, Xinlei He, Yang ZhangCCS 2022 · 被引用 28 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu 等ACL 2025 · 被引用 10 次
