CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
Jingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei He
摘要
Backdoor attacks significantly compromise the security of large language models by triggering them to output specific and controlled content. Currently, triggers for textual backdoor attacks fall into two categories: fixed-token triggers and sentence-pattern triggers. However, the former are typically easy to identify and filter, while the latter, such as syntax and style, do not apply to all original samples and may lead to semantic shifts. In this paper, inspired by cross-lingual (CL) prompts of LLMs in real-world scenarios, we propose a higher-dimensional trigger method at the paragraph level, namely CL-Attack. CL-Attack injects the backdoor by using texts with specific structures that incorporate multiple languages, thereby offering greater stealthiness and universality compared to existing backdoor attack techniques. Extensive experiments on different tasks and model architectures demonstrate that CL-Attack can achieve nearly 100 percents attack success rate with a low poisoning rate in both classification and generation tasks. We also empirically show that CL-Attack is more robust against current major defense methods compared to baseline backdoor attacks. Additionally, in response to CL-Attack, we further develop a new defense called TranslateDefense, which can partially mitigate the impact of CL-Attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety NeuronsXianhui Zhang, Chengyu Xie, Linxia Zhu, Yonghui Yang 等ICML 2026 · 被引用 5 次
- Rounding-Guided Backdoor Injection in Deep Learning Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Wenhai Wang 等NDSS 2026 · 被引用 4 次
- Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web SearchZeren Luo, Zifan Peng, Yule Liu, Zhen Sun 等USENIX Security 2025
- PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-TuningZhen Sun, Tianshuo Cong, Yule Liu, Chenhao Lin 等S&P 2025
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social MediaZhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang 等ACL 2025
它引用的顶会 Paper13
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 被引用 230 次
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao 等CCS 2021 · 被引用 108 次
- Data Poisoning Attacks Against Multimodal EncodersZiqing Yang, Xinlei He, Zheng Li, Michael Backes 等ICML 2023 · 被引用 74 次
相关 Paper
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen 等USENIX Security 2025
- MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language ModelsZihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng 等AAAI 2026
- CROW: Eliminating Backdoors from Large Language Models via Internal Consistency RegularizationNay Myat Min, Long H. Pham, Yige Li, Jun SunICML 2025
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang 等ACL 2021
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou 等ACL 2026 · 被引用 4 次
