Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric Method
Yang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek, Katie Shilton, Hal Daumé III
Abstract
Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether improvements on those tasks have truly addressed moderators' needs in accomplishing their work. In this paper, we surface gaps between past research efforts that have aimed to provide automation for aspects of content moderation and the needs of volunteer content moderators, regarding identifying violations of various moderation rules. To do so, we conduct a model review on Hugging Face to reveal the availability of models to cover various moderation rules and guidelines from three exemplar forums. We further put state-of-the-art LLMs to the test, evaluating how well these models perform in flagging violations of platform rules from one particular forum. Finally, we conduct a user survey study with volunteer moderators to gain insight into their perspectives on useful moderation models. Overall, we observe a nontrivial gap, as missing developed models and LLMs exhibit moderate to low performance on a significant portion of the rules. Moderators' reports provide guides for future work on developing moderation assistant models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 720ec20b-fdfc-40c0-a5c8-1ca06e5ebf9eCited by top-tier papers2
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language ModelsHuy Nghiem, Advik Sachdeva, Hal Daumé IIIACL 2026 · 1 citation
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding et al.EMNLP 2025 · 1 citation
Builds on2
- Hate Raids on Twitch: Echoes of the Past, New Modalities, and Implications for Platform GovernanceCatherine Han, Joseph Seering, Deepak Kumar, Jeffrey T. Hancock et al.CSCW 2023 · 43 citations
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain et al.ACL 2020 · 11 citations
Related papers
- The Unsung Heroes of Facebook Groups Moderation: A Case Study of Moderation Practices and ToolsTina Kuo, Alicia Hernani, Jens GrossklagsCSCW 2023 · 25 citations
- Supporting Human Raters with the Detection of Harmful Content Using Large Language ModelsKurt Thomas, Patrick Gage Kelley, David Tao, Sarah Meiklejohn et al.S&P 2025
- Please note that I'm just an AI: Analysis of Behavior Patterns of LLMs in (Non-)offensive Speech IdentificationEsra Dönmez, Thang Vu, Agnieszka FalenskaEMNLP 2024 · 1 citation
- "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI ProductsLan Gao, Oscar Chen, Rachel Lee, Nick Feamster et al.USENIX Security 2025
- PluRule: A Benchmark for Moderating Pluralistic Communities on Social MediaZoher Kachwala, Bao Tran Truong, Rasika Muralidharan, Haewoon Kwak et al.ACL 2026
