Using Off-the-Shelf Harmful Content Detection Models: Best Practices for Model Reuse
Angela M. Schöpke-Gonzalez, Siqi Wu, Sagar Kumar, Libby Hemphill
Abstract
Supervised machine learning is a common approach for automated harmful content detection to support content moderation. This approach relies on data annotated by humans to train models to recognize classes of harmful content. For detection tasks, researchers or content moderation communities typically either design their own annotation tasks to generate training data for new harmful content detection models, or use off-the-shelf (OTS) pre-trained harmful content detection models. OTS model reuse can enable detection tasks in resource-constrained contexts and can help to reduce the environmental impact of training new models -- an energy-intensive process. However, given the plethora of OTS models now available for reuse, determining which OTS model to reuse for a particular task and how to use it can be challenging, especially given that many of these models have been developed for specific contexts that are not always easily transferred onto others. This work aims to provide best practices for reusing OTS models for harmful content detection tasks. By using content analysis and statistical methods to evaluate assumptions about OTS model utility and reusability, we show that model reusers cannot assume that a model claimed to detect a particular concept, will actually detect that concept. Instead, based on our findings, we offer a decision tree for how to assess whether an OTS model would be appropriate for reuse for a new harmful content detection task. This decision tree directs model reusers to critically assess concept definitions, annotation task design, and additional features specified in our content analysis codebook to identify expected model output, and consequently evaluate whether that OTS model is appropriate for reuse for a new detection task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76264a1b-aab9-40bf-b76a-30a5f8aa2e5eCited by top-tier papers2
- Evaluating LLM-contaminated Crowdsourcing Data Without Ground TruthYichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang LiuNeurIPS 2025 · 3 citations
- Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right OnlineElisabeth Steffen, Helena MihaljevićCSCW 2026
Builds on5
- Don't You Know That You're Toxic: Normalization of Toxicity in Online GamingNicole A. Beres, Julian Frommel, Elizabeth Reid, Regan L. Mandryk et al.CHI 2021 · 235 citations
- The Structure of Toxic Conversations on TwitterMartin Saveski, Brandon Roy, Deb RoyWWW 2021 · 111 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn et al.EMNLP 2022 · 41 citations
- The Online Identity Help Center: Designing and Developing a Content Moderation Policy Resource for Marginalized Social Media UsersSamuel Mayworm, Shannon Li, Hibby Thach, Daniel Delmonaco et al.CSCW 2024 · 10 citations
Related papers
- Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media PlatformsRajvardhan Oak, Muhammad Haroon, Claire Wonjeong Jo, Magdalena Wojcieszak et al.ACL 2025 · 1 citation
- All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept ReproductionZiyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang et al.ACL 2026
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek et al.EMNLP 2024 · 4 citations
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq et al.ACL 2024
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer et al.ICSE 2023 · 62 citations
