Using Off-the-Shelf Harmful Content Detection Models: Best Practices for Model Reuse
Angela M. Schöpke-Gonzalez, Siqi Wu, Sagar Kumar, Libby Hemphill
摘要
Supervised machine learning is a common approach for automated harmful content detection to support content moderation. This approach relies on data annotated by humans to train models to recognize classes of harmful content. For detection tasks, researchers or content moderation communities typically either design their own annotation tasks to generate training data for new harmful content detection models, or use off-the-shelf (OTS) pre-trained harmful content detection models. OTS model reuse can enable detection tasks in resource-constrained contexts and can help to reduce the environmental impact of training new models -- an energy-intensive process. However, given the plethora of OTS models now available for reuse, determining which OTS model to reuse for a particular task and how to use it can be challenging, especially given that many of these models have been developed for specific contexts that are not always easily transferred onto others. This work aims to provide best practices for reusing OTS models for harmful content detection tasks. By using content analysis and statistical methods to evaluate assumptions about OTS model utility and reusability, we show that model reusers cannot assume that a model claimed to detect a particular concept, will actually detect that concept. Instead, based on our findings, we offer a decision tree for how to assess whether an OTS model would be appropriate for reuse for a new harmful content detection task. This decision tree directs model reusers to critically assess concept definitions, annotation task design, and additional features specified in our content analysis codebook to identify expected model output, and consequently evaluate whether that OTS model is appropriate for reuse for a new detection task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Evaluating LLM-contaminated Crowdsourcing Data Without Ground TruthYichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang LiuNeurIPS 2025 · 被引用 3 次
- Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right OnlineElisabeth Steffen, Helena MihaljevićCSCW 2026
它引用的顶会 Paper5
- Don't You Know That You're Toxic: Normalization of Toxicity in Online GamingNicole A. Beres, Julian Frommel, Elizabeth Reid, Regan L. Mandryk 等CHI 2021 · 被引用 235 次
- The Structure of Toxic Conversations on TwitterMartin Saveski, Brandon Roy, Deb RoyWWW 2021 · 被引用 111 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn 等EMNLP 2022 · 被引用 41 次
- The Online Identity Help Center: Designing and Developing a Content Moderation Policy Resource for Marginalized Social Media UsersSamuel Mayworm, Shannon Li, Hibby Thach, Daniel Delmonaco 等CSCW 2024 · 被引用 10 次
相关 Paper
- Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media PlatformsRajvardhan Oak, Muhammad Haroon, Claire Wonjeong Jo, Magdalena Wojcieszak 等ACL 2025 · 被引用 1 次
- All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept ReproductionZiyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang 等ACL 2026
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq 等ACL 2024
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer 等ICSE 2023 · 被引用 62 次
