ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP
Lu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Xuan Chen, Guangyu Shen, Xiangyu Zhang
摘要
Backdoor attacks have emerged as a prominent threat to natural language processing (NLP) models, where the presence of specific triggers in the input can lead poisoned models to misclassify these inputs to predetermined target classes. Current detection mechanisms are limited by their inability to address more covert backdoor strategies, such as style-based attacks. In this work, we propose an innovative test-time poisoned sample detection framework that hinges on the interpretability of model predictions, grounded in the semantic meaning of inputs. We contend that triggers (e.g., infrequent words) are not supposed to fundamentally alter the underlying semantic meanings of poisoned samples as they want to stay stealthy. Based on this observation, we hypothesize that while the model's predictions for paraphrased clean samples should remain stable, predictions for poisoned samples should revert to their true labels upon the mutations applied to triggers during the paraphrasing process. We employ ChatGPT, a state-of-the-art large language model, as our paraphraser and formulate the trigger-removal task as a prompt engineering problem. We adopt fuzzing, a technique commonly used for unearthing software vulnerabilities, to discover optimal paraphrase prompts that can effectively eliminate triggers while concurrently maintaining input semantics. Experiments on 4 types of backdoor attacks, including the subtle style backdoors, and 4 distinct datasets demonstrate that our approach surpasses baseline methods, including STRIP, RAP, and ONION, in precision and recall. Prediction: positive (✓) Prediction: positive (⨯) This film has special effects which for it's time are very impressive. Prediction: positive (✓) This movie is so cool! The things they do with the pictures are amazing. Predictable, cf ambitious attempt that falls short of the mark. Not worth sitting through for the tired contrived ending. Prediction: negative (✓) It's easy to guess what will happen and the ending is boring.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided SearchXuan Chen, Yuzhou Nie, Wenbo Guo, Xiangyu ZhangNeurIPS 2024 · 被引用 68 次
- Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug UnearthingAsmita, Yaroslav Oliinyk, Michael Scott, Ryan Tsang 等USENIX Security 2024 · 被引用 56 次
- Causality Based Front-door Defense Against Backdoor Attack on Language ModelsYiran Liu, Xiaoang Xu, Zhiyi Hou, Yang YuICML 2024 · 被引用 12 次
- Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous DrivingXuan Chen, Shiwei Feng, Zikang Xiong, Shengwei An 等NeurIPS 2025 · 被引用 6 次
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong 等USENIX Security 2026 · 被引用 1 次
它引用的顶会 Paper18
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov 等S&P 2021 · 被引用 381 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li 等EMNLP 2021 · 被引用 114 次
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao 等CCS 2021 · 被引用 108 次
相关 Paper
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang 等ACL 2025
- Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsJiaming He, Wenbo Jiang, Guanyu Hou, Wenshu Fan 等AAAI 2025 · 被引用 9 次
- Hidden Trigger Backdoor Attack on NLP Models via Linguistic Style ManipulationXudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu 等USENIX Security 2022
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu 等NeurIPS 2023 · 被引用 18 次
- WeDef: Weakly Supervised Backdoor Defense for Text ClassificationLesheng Jin, Zihan Wang, Jingbo ShangEMNLP 2022 · 被引用 4 次
