Toxicity Detection for Free
Zhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao, David A. Wagner
Abstract
Current LLMs are generally aligned to follow safety requirements and tend to refuse toxic prompts. However, LLMs can fail to refuse toxic prompts or be overcautious and refuse benign examples. In addition, state-of-the-art toxicity detectors have low TPRs at low FPR, incurring high costs in real-world applications where toxic examples are rare. In this paper, we introduce Moderation Using LLM Introspection (MULI), which detects toxic prompts using the information extracted directly from LLMs themselves. We found we can distinguish between benign and toxic prompts from the distribution of the first response token's logits. Using this idea, we build a robust detector of toxic prompts using a sparse logistic regression model on the first response token logits. Our scheme outperforms SOTA detectors under multiple metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03c127f7-df39-4d05-85bd-fd42b66c3563Cited by top-tier papers2
- Programming Refusal with Conditional Activation SteeringBruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling et al.ICLR 2025
- A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass ClassificationGonzalo Ariel Meyoyan, Luciano Del CorroACL 2026
Builds on4
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- Large Language Models as Tool MakersTianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen et al.ICLR 2024 · 283 citations
Related papers
- Quantifying Large Language Model Attacks Through the Lens of Model CognitionXiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai et al.USENIX Security 2026
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi et al.ASE 2024 · 6 citations
- Efficient LLM Moderation with Multi-Layer Latent PrototypesMaciej Chrabaszcz, Filip Szatkowski, Bartosz Wójcik, Jan Dubiński et al.ICML 2026
- Whispering Experts: Neural Interventions for Toxicity Mitigation in Language ModelsXavier Suau, Pieter Delobelle, Katherine Metcalf, Armand Joulin et al.ICML 2024 · 31 citations
- Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive RefinementYuxiao Lu, Lin Xu, Yang Sun, Wenjun Li et al.ICLR 2026 · 3 citations
