ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AI
Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser
摘要
We present the first English corpus study on abusive language towards three conversational AI systems gathered 'in the wild': an opendomain social bot, a rule-based chatbot, and a task-based system. To account for the complexity of the task, we take a more 'nuanced' approach where our ConvAI dataset reflects fine-grained notions of abuse, as well as views from multiple expert annotators. We find that the distribution of abuse is vastly different compared to other commonly used datasets, with more sexually tinted aggression towards the virtual persona of these systems. Finally, we report results from bench-marking existing models against this data. Unsurprisingly, we find that there is substantial room for improvement with F1 scores below 90%. Warning: This paper contains examples of language that some people may find offensive or upsetting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser 等EMNLP 2023 · 被引用 44 次
- FairPrism: Evaluating Fairness-Related Harms in Text GenerationEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett 等ACL 2023 · 被引用 9 次
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 被引用 7 次
- Confidence-based Ensembling of Perspective-aware ModelsSilvia Casola, Soda Marem Lo, Valerio Basile, Simona Frenda 等EMNLP 2023 · 被引用 2 次
- Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!Philipp Heinisch, Matthias Orlikowski, Julia Romberg, Philipp CimianoEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper4
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto 等CHI 2021 · 被引用 100 次
- "Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp GroupsPunyajoy Saha, Binny Mathew, Kiran Garimella, Animesh MukherjeeWWW 2021 · 被引用 66 次
- Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech DatasetsNedjma Ousidhoum, Yangqiu Song, Dit-Yan YeungEMNLP 2020 · 被引用 22 次
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain 等ACL 2020 · 被引用 11 次
相关 Paper
- Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive LanguageMichael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef RuppenhoferEMNLP 2023 · 被引用 1 次
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi 等ACL 2023 · 被引用 5 次
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan 等CHI 2025 · 被引用 122 次
- Empathy Is All You Need: How a Conversational Agent Should Respond to Verbal AbuseHyojin Chin, Lebogang Wame Molefi, Mun Yong YiCHI 2020 · 被引用 89 次
- AUTALIC: A Dataset for Anti-AUTistic Ableist Language In ContextNaba Rizvi, Harper Strickland, Daniel Gitelman, Alexis Morales Flores 等ACL 2025 · 被引用 5 次
