PreferCare: Preference Dataset Copyright Protection in LLM Alignment by Watermark Injection and Verification
Jian Lou, Chenyang Zhang, Xiaoyu Zhang, Kai Wu
Abstract
With the urgent need to enhance the safety of LLM applications, there has been a growing focus on alignment training algorithms designed to keep large language models (LLMs) behaving in alignment with human values. Alignment training algorithms rely heavily on preference datasets, which are essential for finetuning LLMs to follow human preferences. However, generating and annotating these datasets is often costly and labor-intensive, making it critical to protect their copyright against unauthorized use. In this paper, we propose PreferCare, the first framework tailor-made for preference dataset copyright protection via watermark injection and verification. PreferCare comprises two consecutive stages: injection and verification. In the injection stage, a style transfer-based watermark signal and a bi-level watermark optimization process are designed to embed the watermark into the preference dataset. In the verification stage, we employ statistical tests to determine whether a suspect LLM has used the watermarked preference dataset without authorization. Extensive experiments on multiple popular LLMs have demonstrated that PreferCare achieves effectiveness, harmlessness, transferability, and robustness across diverse settings, and can successfully verify the watermark within 20 queries.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 136f413e-4076-4a8f-9d71-5a234eb80c97Related papers
- PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationHongwei Yao, Jian Lou, Zhan Qin, Kui RenS&P 2024 · 43 citations
- Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?Michael-Andrei Panaitescu-Liess, Zora Che, Bang An, Yuancheng Xu et al.AAAI 2025 · 21 citations
- ReMoDetect: Reward Models Recognize Aligned LLM's GenerationsHyunseok Lee, Jihoon Tack, Jinwoo ShinNeurIPS 2024 · 13 citations
- STAMP Your Content: Proving Dataset Membership via Watermarked RephrasingsSaksham Rastogi, Pratyush Maini, Danish PruthiICML 2025
- Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor WatermarkWenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu et al.ACL 2023 · 39 citations
