CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
Jingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei He
Abstract
Backdoor attacks significantly compromise the security of large language models by triggering them to output specific and controlled content. Currently, triggers for textual backdoor attacks fall into two categories: fixed-token triggers and sentence-pattern triggers. However, the former are typically easy to identify and filter, while the latter, such as syntax and style, do not apply to all original samples and may lead to semantic shifts. In this paper, inspired by cross-lingual (CL) prompts of LLMs in real-world scenarios, we propose a higher-dimensional trigger method at the paragraph level, namely CL-Attack. CL-Attack injects the backdoor by using texts with specific structures that incorporate multiple languages, thereby offering greater stealthiness and universality compared to existing backdoor attack techniques. Extensive experiments on different tasks and model architectures demonstrate that CL-Attack can achieve nearly 100 percents attack success rate with a low poisoning rate in both classification and generation tasks. We also empirically show that CL-Attack is more robust against current major defense methods compared to baseline backdoor attacks. Additionally, in response to CL-Attack, we further develop a new defense called TranslateDefense, which can partially mitigate the impact of CL-Attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety NeuronsXianhui Zhang, Chengyu Xie, Linxia Zhu, Yonghui Yang et al.ICML 2026 · 5 citations
- Rounding-Guided Backdoor Injection in Deep Learning Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Wenhai Wang et al.NDSS 2026 · 4 citations
- Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web SearchZeren Luo, Zifan Peng, Yule Liu, Zhen Sun et al.USENIX Security 2025
- PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-TuningZhen Sun, Tianshuo Cong, Yule Liu, Chenhao Lin et al.S&P 2025
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social MediaZhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang et al.ACL 2025
Builds on13
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao et al.CCS 2021 · 108 citations
- Data Poisoning Attacks Against Multimodal EncodersZiqing Yang, Xinlei He, Zheng Li, Michael Backes et al.ICML 2023 · 74 citations
Related papers
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen et al.USENIX Security 2025
- MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language ModelsZihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng et al.AAAI 2026
- CROW: Eliminating Backdoors from Large Language Models via Internal Consistency RegularizationNay Myat Min, Long H. Pham, Yige Li, Jun SunICML 2025
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang et al.ACL 2021
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou et al.ACL 2026 · 4 citations
