Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them
Emanuele Moscato, Tiancheng Hu, Matthias Orlikowski, Paul Röttger, Debora Nozza
摘要
Personalized content moderation can protect users from harm while facilitating free expression by tailoring moderation decisions to individual preferences rather than enforcing universal rules. However, content moderation that is fully personalized to individual preferences, no matter what these preferences are, may lead to even the most hazardous types of content being propagated on social media. In this paper, we explore this risk using hate speech as a case study. Certain types of hate speech are illegal in many countries. We show that, while fully personalized hate speech detection models increase overall user welfare (as measured by user-level classification performance), they also make predictions that violate such legal hate speech boundaries, especially when tailored to users who tolerate highly hateful content. To address this problem, we enforce legal boundaries in personalized hate speech detection by overriding predictions from personalized models with those from a boundary classifier. This approach significantly reduces legal violations while minimally affecting overall user welfare. Our findings highlight both the promise and the risks of personalized moderation, and offer a practical solution to balance user preferences with legal and ethical obligations. Content warning: This paper contains obfuscated examples of hate speech.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and LaborShagun Jhaver, Alice Qian Zhang, Quan Ze Chen, Nikhila Natarajan 等CSCW 2023 · 被引用 87 次
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 被引用 74 次
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano 等ACL 2025 · 被引用 34 次
相关 Paper
- "Ignorance is not Bliss": Designing Personalized Moderation to Address Ableist Hate on Social MediaSharon Heung, Lucy Jiang, Shiri Azenkot, Aditya VashisthaCHI 2025 · 被引用 14 次
- Controversy and Conformity: from Generalized to Personalized Aggressiveness DetectionKamil Kanclerz, Alicja Figas, Marcin Gruza, Tomasz Kajdanowicz 等ACL 2021
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on InstagramAnna Ricarda Luther, Hendrik Heuer, Sebastian Haunss, Stephanie Geise 等CHI 2026 · 被引用 1 次
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
