Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
Yixiao Xu, Binxing Fang, Rui Wang, Yinghai Zhou, Yuan Liu, Mohan Li, Zhihong Tian
摘要
Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post-deployment flexibility, and the lack of clear theoretical foundations makes them vulnerable to adaptive attacks. In this paper, we propose Neural Honeytrace, a plug-and-play watermarking framework that operates without retraining. We redefine the watermark transmission mechanism from an information perspective, designing a training-free multi-step transmission strategy that leverages the long-tailed effect of backdoor learning to achieve efficient and robust watermark embedding. Extensive experiments demonstrate that Neural Honeytrace reduces the average number of queries required for a worst-case t-test-based ownership verification to as low as 2% of existing methods, while incurring zero training cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 被引用 287 次
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 被引用 197 次
- SoK: How Robust is Image Classification Deep Neural Network Watermarking?Nils Lukas, Edward Jiang, Xinda Li, Florian KerschbaumS&P 2022 · 被引用 124 次
- Membership Inference Attacks by Exploiting Loss TrajectoryYiyong Liu, Zhengyu Zhao, Michael Backes, Yang ZhangCCS 2022 · 被引用 79 次
相关 Paper
- Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature AttributionShuo Shao, Yiming Li, Hongwei Yao, Yiling He 等NDSS 2025
- Safe and Robust Watermark Injection with a Single OoD ImageShuyang Yu, Junyuan Hong, Haobo Zhang, Haotao Wang 等ICLR 2024 · 被引用 4 次
- Watermarking Graph Neural Networks via Explanations for Ownership ProtectionJane Downer, Yingdan Shi, Ziyan Liu, Ren Wang 等ICML 2026 · 被引用 3 次
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 被引用 18 次
- Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model WatermarkingCong Kong, Rui Xu, Jiawei Chen, Zhaoxia YinACM MM 2025 · 被引用 1 次
