HateCheck: Functional Tests for Hate Speech Detection Models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Z. Margetts, Janet B. Pierrehumbert
摘要
Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HATECHECK, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HATECHECK's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain 等NeurIPS 2021 · 被引用 76 次
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 被引用 74 次
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 被引用 74 次
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith 等EMNLP 2022 · 被引用 54 次
它引用的顶会 Paper8
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 被引用 93 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
- The Curse of Performance Instability in Analysis Datasets: Consequences, Source, and SuggestionsXiang Zhou, Yixin Nie, Hao Tan, Mohit BansalEMNLP 2020 · 被引用 30 次
相关 Paper
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes 等USENIX Security 2025
- Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech DatasetsNedjma Ousidhoum, Yangqiu Song, Dit-Yan YeungEMNLP 2020 · 被引用 22 次
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 被引用 11 次
