CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
Rui Zeng, Xi Chen, Yuwen Pu, Xuhong Zhang, Tianyu Du, Shouling Ji
Abstract
Backdoors can be injected into NLP models to induce misbehavior when the input text contains a specific feature, known as a trigger, which the attacker secretly selects. Unlike fixed tokens, words, phrases, or sentences used in the textitstatic text trigger, textitdynamic backdoor attacks on NLP models design triggers associated with abstract and latent text features (e.g., style), making them considerably stealthier than traditional static backdoor attacks. However, existing research on NLP backdoor detection primarily focuses on defending against static backdoor attacks, while research on detecting dynamic backdoors in NLP models remains largely unexplored.
This paper presents CLIBE, the first framework to detect dynamic backdoors in Transformer-based NLP models. At a high level, CLIBE injects a textit"few-shot perturbation" into the suspect Transformer model by crafting an optimized weight perturbation in the attention layers to make the perturbed model classify a limited number of reference samples as a target label. Subsequently, CLIBE leverages the textitgeneralization capability of this "few-shot perturbation" to determine whether the original suspect model contains a dynamic backdoor. Extensive evaluation on three advanced NLP dynamic backdoor attacks, two widely-used Transformer frameworks, and four real-world classification tasks strongly validates the effectiveness and generality of CLIBE. We also demonstrate the robustness of CLIBE against various adaptive attacks. Furthermore, we employ CLIBE to scrutinize 49 popular Transformer models on Hugging Face and discover one model exhibiting a high probability of containing a dynamic backdoor. We have contacted Hugging Face and provided detailed evidence of the backdoor behavior of this model. Moreover, we show that CLIBE can be easily extended to detect backdoor text generation models (e.g., GPT-Neo-1.3B) that are modified to exhibit toxic behavior. To the best of our knowledge, CLIBE is the first framework capable of detecting backdoors in text generation models without requiring access to trigger input test samples. The code is available at https://github.com/Raytsang123/CLIBE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c20f4812-8fb8-4bc6-bf53-e783cf640af6Cited by top-tier papers6
- ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context IlluminationXiaoyi Pang, Xuanyi Hao, Song Guo, Qi Luo et al.NeurIPS 2025 · 7 citations
- Improving the Sensitivity of Backdoor Detectors via Class Subspace OrthogonalizationGuangmingmei Yang, David Miller, George KesidisICML 2026 · 3 citations
- Tell me about yourself: LLMs are aware of their learned behaviorsJan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley et al.ICLR 2025 · 2 citations
- TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image ModelsXindi Li, Zhe Liu, Tong Zhang, Jiahao Chen et al.ACL 2025 · 1 citation
- Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement LearningOubo Ma, Ruixiao Lin, Yang Dai, Jiahao Chen et al.ICML 2026 · 1 citation
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
Related papers
- Piccolo: Exposing Complex Backdoors in NLP Transformer ModelsYingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An et al.S&P 2022 · 100 citations
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li et al.CCS 2021 · 72 citations
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang et al.ACL 2021
- Hidden Trigger Backdoor Attack on NLP Models via Linguistic Style ManipulationXudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu et al.USENIX Security 2022
- Constrained Optimization with Dynamic Bound-scaling for Effective NLP Backdoor DefenseGuangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu et al.ICML 2022 · 58 citations
