ModelCitizens: Representing Community Voices in Online Safety
Ashima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi, Sophie Hao, Thomas Hartvigsen, Saadia Gabriel
摘要
Warning: This paper contains content that may be offensive or upsetting. Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity detection models are typically trained on annotations that collapse diverse annotator perspectives into a single ground truth, erasing important context-specific notions of toxicity such as reclaimed language. To address this, we introduce MODELCITIZENS, a dataset of 6.8K social media posts and 40K toxicity annotations across diverse identity groups. To capture the role of conversational context on toxicity, typical of social media posts, we augment MODEL-CITIZENS posts with LLM-generated conversational scenarios. State-of-the-art toxicity detection tools (e.g. OpenAI Moderation API, GPT-o4-mini) underperform on MODELCITIZENS, with further degradation on context-augmented posts. Finally, we release LLAMACITIZEN-8B and GEMMACITIZEN-12B, LLaMA-and Gemma-based models finetuned on MODEL-CITIZENS, which outperform GPT-o4-mini by 5.5% on in-distribution evaluations. Our findings highlight the importance of communityinformed annotation and modeling for inclusive content moderation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 被引用 74 次
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke 等ACL 2023 · 被引用 23 次
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky 等ACL 2020 · 被引用 16 次
- Beyond Initial Removal: Lasting Impacts of Discriminatory Content Moderation to Marginalized Creators on InstagramYim Register, Izzi Grasso, Lauren N. Weingarten, Lilith Fury 等CSCW 2024 · 被引用 14 次
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain 等ACL 2020 · 被引用 11 次
相关 Paper
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 等ACL 2022
- UniDetox: Universal Detoxification of Large Language Models via Dataset DistillationHuimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro SakataICLR 2025
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi 等ASE 2024 · 被引用 6 次
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding 等EMNLP 2025 · 被引用 1 次
- Toxicity Detection for FreeZhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao 等NeurIPS 2024 · 被引用 20 次
