TrojText: Test-time Invisible Textual Trojan Insertion
Qian Lou, Yepeng Liu, Bo Feng
摘要
In Natural Language Processing (NLP), intelligent neuron models can be susceptible to textual Trojan attacks. Such attacks occur when Trojan models behave normally for standard inputs but generate malicious output for inputs that contain a specific trigger. Syntactic-structure triggers, which are invisible, are becoming more popular for Trojan attacks because they are difficult to detect and defend against. However, these types of attacks require a large corpus of training data to generate poisoned samples with the necessary syntactic structures for Trojan insertion. Obtaining such data can be difficult for attackers, and the process of generating syntactic poisoned triggers and inserting Trojans can be time-consuming. This paper proposes a solution called TrojText, which aims to determine whether invisible textual Trojan attacks can be performed more efficiently and cost-effectively without training data. The proposed approach, called the Representation-Logit Trojan Insertion (RLI) algorithm, uses smaller sampled test data instead of large training data to achieve the desired attack. The paper also introduces two additional techniques, namely the accumulated gradient ranking (AGR) and Trojan Weights Pruning (TWP), to reduce the number of tuned parameters and the attack overhead. The TrojText approach was evaluated on three datasets (AG's News, SST-2, and OLID) using three NLP models (BERT, XL-Net, and DeBERTa). The experiments demonstrated that the TrojText approach achieved a 98.35% classification accuracy for test sentences in the target class on the BERT model for the AG's News dataset. The source code for TrojText is available at https://github.com/UCF-ML-Research/TrojText .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Memory Injection Attacks on LLM Agents via Query-Only InteractionShen Dong, Shaochen Xu, Pengfei He, Yige Li 等NeurIPS 2025 · 被引用 146 次
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian 等ICLR 2024 · 被引用 98 次
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language ModelsJiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen 等NeurIPS 2023 · 被引用 63 次
- Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context LearningShuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan 等EMNLP 2024 · 被引用 28 次
- Jailbreaking LLMs with Arabic Transliteration and ArabiziMansour Al Ghanim, Saleh Almohaimeed, Mengxin Zheng, Yan Solihin 等EMNLP 2024 · 被引用 3 次
它引用的顶会 Paper14
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- CSI NN: Reverse Engineering of Neural Network Architectures Through Electromagnetic Side ChannelLejla Batina, Shivam Bhasin, Dirmanto Jap, Stjepan PicekUSENIX Security 2019 · 被引用 334 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Bit-Flip Attack: Crushing Neural Network With Progressive Bit SearchAdnan Siraj Rakin, Zhezhi He, Deliang FanICCV 2019 · 被引用 309 次
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou 等ICLR 2022 · 被引用 210 次
相关 Paper
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang 等ACL 2021
- T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text ClassificationAhmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar 等USENIX Security 2021 · 被引用 98 次
- TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image ModelsXindi Li, Zhe Liu, Tong Zhang, Jiahao Chen 等ACL 2025 · 被引用 1 次
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao 等CCS 2021 · 被引用 108 次
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang 等KDD 2020 · 被引用 164 次
