ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AI
Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser
Abstract
We present the first English corpus study on abusive language towards three conversational AI systems gathered 'in the wild': an opendomain social bot, a rule-based chatbot, and a task-based system. To account for the complexity of the task, we take a more 'nuanced' approach where our ConvAI dataset reflects fine-grained notions of abuse, as well as views from multiple expert annotators. We find that the distribution of abuse is vastly different compared to other commonly used datasets, with more sexually tinted aggression towards the virtual persona of these systems. Finally, we report results from bench-marking existing models against this data. Unsurprisingly, we find that there is substantial room for improvement with F1 scores below 90%. Warning: This paper contains examples of language that some people may find offensive or upsetting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52e4de68-808c-47a1-8e90-443ce3df70b6Cited by top-tier papers5
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser et al.EMNLP 2023 · 44 citations
- FairPrism: Evaluating Fairness-Related Harms in Text GenerationEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett et al.ACL 2023 · 9 citations
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 7 citations
- Confidence-based Ensembling of Perspective-aware ModelsSilvia Casola, Soda Marem Lo, Valerio Basile, Simona Frenda et al.EMNLP 2023 · 2 citations
- Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!Philipp Heinisch, Matthias Orlikowski, Julia Romberg, Philipp CimianoEMNLP 2023 · 1 citation
Builds on4
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto et al.CHI 2021 · 100 citations
- "Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp GroupsPunyajoy Saha, Binny Mathew, Kiran Garimella, Animesh MukherjeeWWW 2021 · 66 citations
- Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech DatasetsNedjma Ousidhoum, Yangqiu Song, Dit-Yan YeungEMNLP 2020 · 22 citations
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain et al.ACL 2020 · 11 citations
Related papers
- Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive LanguageMichael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef RuppenhoferEMNLP 2023 · 1 citation
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi et al.ACL 2023 · 5 citations
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan et al.CHI 2025 · 122 citations
- Empathy Is All You Need: How a Conversational Agent Should Respond to Verbal AbuseHyojin Chin, Lebogang Wame Molefi, Mun Yong YiCHI 2020 · 89 citations
- AUTALIC: A Dataset for Anti-AUTistic Ableist Language In ContextNaba Rizvi, Harper Strickland, Daniel Gitelman, Alexis Morales Flores et al.ACL 2025 · 5 citations
