Efficient Detection of Toxic Prompts in Large Language Models
Yi Liu, Junzhe Yu, Huijia Sun, Ling Shi, Gelei Deng, Yuqi Chen, Yang Liu
Abstract
Large language models (LLMs) like ChatGPT and Gemini have significantly advanced natural language processing, enabling various applications such as chatbots and automated content generation. However, these models can be exploited by malicious individuals who craft toxic prompts to elicit harmful or unethical responses. These individuals often employ jailbreaking techniques to bypass safety mechanisms, highlighting the need for robust toxic prompt detection methods. Existing detection techniques, both blackbox and whitebox, face challenges related to the diversity of toxic prompts, scalability, and computational efficiency. In response, we propose ToxicDetector, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs. ToxicDetector leverages LLMs to create toxic concept prompts, uses embedding vectors to form feature vectors, and employs a Multi-Layer Perceptron (MLP) classifier for prompt classification. Our evaluation on various versions of the LLama models, Gemma-2, and multiple datasets demonstrates that ToxicDetector achieves a high accuracy of 96.39% and a low false positive rate of 2.00%, outperforming state-of-the-art methods. Additionally, ToxicDetector's processing time of 0.0780 seconds per prompt makes it highly suitable for real-time applications. ToxicDetector achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac4d4a36-066a-4cdd-bf9a-bcd5bc674004Cited by top-tier papers5
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive ScoringPeichun Hua, Hao Li, Shanghao Shi, Zhiyuan Yu et al.ACL 2026 · 8 citations
- Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory KineticsHangtao Zhang, Yucheng Zhao, Sishun Liu, Ziqi Zhou et al.USENIX Security 2026 · 3 citations
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM AgentsHaoyu Wang, Christopher M. Poskitt, Jun SunICSE 2026 · 2 citations
- Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier MonitoringGuanxu Chen, Jing Shao, Tao Luo, Lijie Hu et al.ICML 2026 · 2 citations
- A Strategic Coordination Framework of Small LMs Matches Large LMs in Data SynthesisXin Gao, Qizhi Pei, Zinan Tang, Yu Li et al.ACL 2025
Builds on5
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 836 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain ChatbotsWai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro et al.CCS 2022 · 34 citations
- Defending Jailbreak Prompts via In-Context Adversarial GameYujun Zhou, Yufei Han, Haomin Zhuang, Kehan Guo et al.EMNLP 2024 · 9 citations
Related papers
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi Xie, Minghong Fang, Renjie Pi, Neil GongACL 2024
- On Large Language Models' Resilience to Coercive InterrogationZhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng et al.S&P 2024 · 24 citations
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 74 citations
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu et al.USENIX Security 2025
