ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Ece Kamar
Abstract
Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language. To help mitigate these issues, we create TOXIGEN, a new large-scale and machinegenerated dataset of 274k toxic and benign statements about 13 minority groups. We develop a demonstration-based prompting framework and an adversarial classifier-in-the-loop decoding method to generate subtly toxic and benign text with a massive pretrained language model (Brown et al., 2020) . Controlling machine generation in this way allows TOXIGEN to cover implicitly toxic text at a larger scale, and about more demographic groups, than previous resources of human-written text. We conduct a human evaluation on a challenging subset of TOXIGEN and find that annotators struggle to distinguish machine-generated text from human-written language. We also find that 94.5% of toxic examples are labeled as hate speech by human annotators. Using three publicly-available datasets, we show that finetuning a toxicity classifier on our data improves its performance on human-written data substantially. We also demonstrate that TOXI-GEN can be used to fight machine-generated toxicity as finetuning improves the classifier significantly on our evaluation subset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75f7a475-3a17-4409-882a-b6bf8973069cCited by top-tier papers156
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- Design Principles for Generative AI ApplicationsJustin D. Weisz, Jessica He, Michael J. Muller, Gabriela Hoefer et al.CHI 2024 · 221 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsWenlong Huang, Fei Xia, Dhruv Shah, Danny Driess et al.NeurIPS 2023 · 102 citations
Builds on7
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem et al.ACL 2021
Related papers
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi et al.ASE 2024 · 6 citations
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi et al.EMNLP 2025
- When Bad Data Leads to Good ModelsKenneth Li, Yida Chen, Fernanda B. Viégas, Martin WattenbergICML 2025
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu et al.NeurIPS 2022 · 89 citations
