Antidistillation Fingerprinting
Yixuan Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein, Fei Fang, Zico Kolter
摘要
Model distillation enables efficient emulation of frontier large language models (LLMs), creating a need for robust mechanisms to detect when a third-party student model has trained on a teacher model's outputs. However, existing fingerprinting techniques that could be used to detect such distillation rely on heuristic perturbations that impose a steep trade-off between generation quality and fingerprinting strength, often requiring significant degradation of utility to ensure the fingerprint is effectively internalized by the student. We introduce antidistillation fingerprinting (ADFP), a principled approach that aligns the fingerprinting objective with the student's learning dynamics. Building upon the gradient-based framework of antidistillation sampling , ADFP utilizes a proxy model to identify and sample tokens that directly maximize the expected detectability of the fingerprint in the student after fine-tuning, rather than relying on the incidental absorption of the un-targeted biases of a more naive watermark. Experiments on GSM8K, OASST1, and MBPP demonstrate that ADFP achieves a significant Pareto improvement over state-of-the-art baselines, yielding stronger detection confidence with minimal impact on utility across mathematical reasoning, dialogue, and code generation, even when the student model's architecture is unknown.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
相关 Paper
- Protecting Language Models Against Unauthorized Distillation through Trace RewritingXinhang Ma, William Yeoh, Ning Zhang, Yevgeniy VorobeychikACL 2026 · 被引用 6 次
- Antidistillation SamplingYash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu 等NeurIPS 2025 · 被引用 24 次
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu 等ACL 2025 · 被引用 10 次
- Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional FingerprintsZirui Huang, Yunlong Mao, Wei Tong, Tingting Wu 等CCS 2026
- ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation AttacksPeizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang 等ACL 2026
