ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP
Lu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Xuan Chen, Guangyu Shen, Xiangyu Zhang
Abstract
Backdoor attacks have emerged as a prominent threat to natural language processing (NLP) models, where the presence of specific triggers in the input can lead poisoned models to misclassify these inputs to predetermined target classes. Current detection mechanisms are limited by their inability to address more covert backdoor strategies, such as style-based attacks. In this work, we propose an innovative test-time poisoned sample detection framework that hinges on the interpretability of model predictions, grounded in the semantic meaning of inputs. We contend that triggers (e.g., infrequent words) are not supposed to fundamentally alter the underlying semantic meanings of poisoned samples as they want to stay stealthy. Based on this observation, we hypothesize that while the model's predictions for paraphrased clean samples should remain stable, predictions for poisoned samples should revert to their true labels upon the mutations applied to triggers during the paraphrasing process. We employ ChatGPT, a state-of-the-art large language model, as our paraphraser and formulate the trigger-removal task as a prompt engineering problem. We adopt fuzzing, a technique commonly used for unearthing software vulnerabilities, to discover optimal paraphrase prompts that can effectively eliminate triggers while concurrently maintaining input semantics. Experiments on 4 types of backdoor attacks, including the subtle style backdoors, and 4 distinct datasets demonstrate that our approach surpasses baseline methods, including STRIP, RAP, and ONION, in precision and recall. Prediction: positive (✓) Prediction: positive (⨯) This film has special effects which for it's time are very impressive. Prediction: positive (✓) This movie is so cool! The things they do with the pictures are amazing. Predictable, cf ambitious attempt that falls short of the mark. Not worth sitting through for the tired contrived ending. Prediction: negative (✓) It's easy to guess what will happen and the ending is boring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dde47d7f-90cb-42e1-b908-6f3694e96487Cited by top-tier papers8
- When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided SearchXuan Chen, Yuzhou Nie, Wenbo Guo, Xiangyu ZhangNeurIPS 2024 · 68 citations
- Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug UnearthingAsmita, Yaroslav Oliinyk, Michael Scott, Ryan Tsang et al.USENIX Security 2024 · 56 citations
- Causality Based Front-door Defense Against Backdoor Attack on Language ModelsYiran Liu, Xiaoang Xu, Zhiyi Hou, Yang YuICML 2024 · 12 citations
- Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous DrivingXuan Chen, Shiwei Feng, Zikang Xiong, Shengwei An et al.NeurIPS 2025 · 6 citations
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong et al.USENIX Security 2026 · 1 citation
Builds on18
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov et al.S&P 2021 · 381 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao et al.CCS 2021 · 108 citations
Related papers
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang et al.ACL 2025
- Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsJiaming He, Wenbo Jiang, Guanyu Hou, Wenshu Fan et al.AAAI 2025 · 9 citations
- Hidden Trigger Backdoor Attack on NLP Models via Linguistic Style ManipulationXudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu et al.USENIX Security 2022
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu et al.NeurIPS 2023 · 18 citations
- WeDef: Weakly Supervised Backdoor Defense for Text ClassificationLesheng Jin, Zihan Wang, Jingbo ShangEMNLP 2022 · 4 citations
