USENIX Security2023Top-tier venue
Two-in-One: A Model Hijacking Attack Against Text Generation Models
Wai Man Si, Michael Backes, Yang Zhang, Ahmed Salem
Abstract
Machine learning has progressed significantly in various applications ranging from face recognition to text generation. However, its success has been accompanied by different attacks. Recently a new attack has been proposed which raises both accountability and parasitic computing risks, namely the model hijacking attack. Nevertheless, this attack has only focused on image classification tasks. In this work, we broaden the scope of this attack to include text generation and classification models, hence showing its broader applicability. More concretely, we propose a new model hijacking attack, Ditto, that can hijack different text classification tasks into multiple generation ones, e.g., language translation, text summarization, and language modeling. We use a range of text benchmark datasets such as SST-2, TweetEval, AGnews, QNLI, and IMDB to evaluate the performance of our attacks. Our results show that by using Ditto, an adversary can successfully hijack text generation models without jeopardizing their utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88b4f34f-4250-403a-a053-05bb807b537fCited by top-tier papers5
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial ManipulationRui Chu, Bingyin Zhao, Hanling Jiang, Shuchin Aeron et al.NeurIPS 2025 · 4 citations
- Mithridates: Auditing and Boosting Backdoor Resistance of Machine Learning PipelinesEugene Bagdasarian, Vitaly ShmatikovCCS 2024 · 3 citations
- CAMH: Advancing Model Hijacking Attack in Machine LearningXing He, Jiahao Chen, Yuwen Pu, Qingming Li et al.AAAI 2025 · 1 citation
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang et al.NDSS 2024
Builds on15
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
Related papers
- Get a Model! Model Hijacking Attack Against Machine Learning ModelsAhmed Salem, Michael Backes, Yang ZhangNDSS 2022
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language ModelsZenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou et al.ACL 2025 · 5 citations
- Phi: Preference Hijacking in Multi-modal Large Language Models at Inference TimeYifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin et al.EMNLP 2025
- Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingJunxi Chen, Junhao Dong, Xiaohua XieCVPR 2025
- Data-Free Model-Related Attacks: Unleashing the Potential of Generative AIDayong Ye, Tianqing Zhu, Shang Wang, Bo Liu et al.USENIX Security 2025
