Jailbreaking the Non-Transferable Barrier via Test-Time Data Disguising
Yongli Xiang, Ziming Hong, Lina Yao, Dadong Wang, Tongliang Liu
摘要
Non-transferable learning (NTL) has been proposed to protect model intellectual property (IP) by creating a "nontransferable barrier" to restrict generalization from authorized to unauthorized domains. Recently, well-designed attack, which restores the unauthorized-domain performance by fine-tuning NTL models on few authorized samples, highlights the security risks of NTL-based applications. However, such attack requires modifying model weights, thus being invalid in the black-box scenario. This raises a critical question: can we trust the security of NTL models deployed as black-box systems? In this work, we reveal the first loophole of black-box NTL models by proposing a novel attack method (dubbed as JailNTL) to jailbreak the non-transferable barrier through test-time data disguising. The main idea of JailNTL is to disguise unauthorized data so it can be identified as authorized by the NTL model, thereby bypassing the non-transferable barrier without modifying the NTL model weights. Specifically, JailNTL encourages unauthorized-domain disguising in two levels, including: (i) data-intrinsic disguising (DID) for eliminating domain discrepancy and preserving class-related content at the input-level, and (ii) model-guided disguising (MGD) for mitigating output-level statistics difference of the NTL model. Empirically, when attacking state-of-theart (SOTA) NTL models in the black-box scenario, Jail-NTL achieves an accuracy increase of up to 55.7% in the unauthorized domain by using only 1% authorized samples, largely exceeding existing SOTA white-box attacks. Code is released at https://github.com/tmllab/2025_ CVPR_JailNTL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety GuidanceYongli Xiang, Ziming Hong, Zhaoqing Wang, Xiangyu Zhao 等CVPR 2026 · 被引用 14 次
- AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven EditingZiming Hong, Tianyu Huang, Runnan Chen, Shanshan Ye 等ICML 2026 · 被引用 10 次
- When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You NeedZiming Hong, Runnan Chen, Zengmao Wang, Bo Han 等ICML 2025
它引用的顶会 Paper23
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen 等ICML 2022 · 被引用 579 次
- Test-Time Classifier Adjustment Module for Model-Agnostic Domain GeneralizationYusuke Iwasawa, Yutaka MatsuoNeurIPS 2021 · 被引用 456 次
相关 Paper
- Your Transferability Barrier is Fragile: Free-Lunch for Transferring the Non-Transferable LearningZiming Hong, Li Shen, Tongliang LiuCVPR 2024
- Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability AuthorizationLixu Wang, Shichao Xu, Ruiqi Xu, Xiao Wang 等ICLR 2022 · 被引用 65 次
- Improving Non-Transferable Representation Learning by Harnessing Content and StyleZiming Hong, Zhenyi Wang, Li Shen, Yu Yao 等ICLR 2024 · 被引用 37 次
- Model Barrier: A Compact Un-Transferable Isolation Domain for Model Intellectual Property ProtectionLianyu Wang, Meng Wang, Daoqiang Zhang, Huazhu FuCVPR 2023
- Unsupervised Non-transferable Text ClassificationGuangtao Zeng, Wei LuEMNLP 2022 · 被引用 4 次
