Two-in-One: A Model Hijacking Attack Against Text Generation Models
Wai Man Si, Michael Backes, Yang Zhang, Ahmed Salem
摘要
Machine learning has progressed significantly in various applications ranging from face recognition to text generation. However, its success has been accompanied by different attacks. Recently a new attack has been proposed which raises both accountability and parasitic computing risks, namely the model hijacking attack. Nevertheless, this attack has only focused on image classification tasks. In this work, we broaden the scope of this attack to include text generation and classification models, hence showing its broader applicability. More concretely, we propose a new model hijacking attack, Ditto, that can hijack different text classification tasks into multiple generation ones, e.g., language translation, text summarization, and language modeling. We use a range of text benchmark datasets such as SST-2, TweetEval, AGnews, QNLI, and IMDB to evaluate the performance of our attacks. Our results show that by using Ditto, an adversary can successfully hijack text generation models without jeopardizing their utility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial ManipulationRui Chu, Bingyin Zhao, Hanling Jiang, Shuchin Aeron 等NeurIPS 2025 · 被引用 4 次
- Mithridates: Auditing and Boosting Backdoor Resistance of Machine Learning PipelinesEugene Bagdasarian, Vitaly ShmatikovCCS 2024 · 被引用 3 次
- CAMH: Advancing Model Hijacking Attack in Machine LearningXing He, Jiahao Chen, Yuwen Pu, Qingming Li 等AAAI 2025 · 被引用 1 次
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang 等NDSS 2024
它引用的顶会 Paper15
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu 等S&P 2018 · 被引用 867 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
相关 Paper
- Get a Model! Model Hijacking Attack Against Machine Learning ModelsAhmed Salem, Michael Backes, Yang ZhangNDSS 2022
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language ModelsZenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou 等ACL 2025 · 被引用 5 次
- Phi: Preference Hijacking in Multi-modal Large Language Models at Inference TimeYifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin 等EMNLP 2025
- Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingJunxi Chen, Junhao Dong, Xiaohua XieCVPR 2025
- Data-Free Model-Related Attacks: Unleashing the Potential of Generative AIDayong Ye, Tianqing Zhu, Shang Wang, Bo Liu 等USENIX Security 2025
