From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer
John Pavlopoulos, Léo Laugier, Alexandros Xenos, Jeffrey Sorensen, Ion Androutsopoulos
Abstract
We study the task of toxic spans detection, which concerns the detection of the spans that make a text toxic, when detecting such spans is possible. We introduce a dataset for this task, TOXICSPANS, which we release publicly. By experimenting with several methods, we show that sequence labeling models perform best. Moreover, methods that add generic rationale extraction mechanisms on top of classifiers trained to predict if a post is toxic or not are also surprisingly promising. Finally, we use TOXICSPANS and systems trained on it, to provide further analysis of state-of-the-art toxic to non-toxic transfer systems, as well as of human performance on that latter task. Our work highlights challenges in finer toxicity detection and mitigation. Gold Spans (set of character offsets)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2d69469-e37b-4269-b2c3-863cd24967d0Cited by top-tier papers4
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 74 citations
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn et al.EMNLP 2022 · 41 citations
- Eyes Don't Lie: Subjective Hate Annotation and Detection with GazeÖzge Alaçam, Sanne Hoeken, Sina ZarrießEMNLP 2024 · 1 citation
- UnGANable: Defending Against GAN-based Face ManipulationZheng Li, Ning Yu, Ahmed Salem, Michael Backes et al.USENIX Security 2023
Builds on5
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain et al.ACL 2020 · 11 citations
- Learning to Faithfully Rationalize by ConstructionSarthak Jain, Sarah Wiegreffe, Yuval Pinter, Byron C. WallaceACL 2020
Related papers
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi et al.ACL 2023 · 5 citations
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap et al.ACL 2022
- Exploring Modular Task Decomposition in Cross-domain Named Entity RecognitionXinghua Zhang, Bowen Yu, Yubin Wang, Tingwen Liu et al.SIGIR 2022 · 18 citations
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy et al.ACL 2022 · 96 citations
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding et al.EMNLP 2025 · 1 citation
