Supporting Human Raters with the Detection of Harmful Content Using Large Language Models
Kurt Thomas, Patrick Gage Kelley, David Tao, Sarah Meiklejohn, Owen Vallis, Shunwen Tan, Blaz Bratanic, Felipe Tiengo Ferreira, Vijay Kumar Eranti, Elie Bursztein
Abstract
In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election misinformation. Using a dataset of 50,000 user comments, we demonstrate that LLMs can achieve 90 % accuracy when compared to human verdicts. We explore how to best leverage these capabilities, proposing five design patterns that integrate LLMs with human rating, such as pre-filtering non-violative content, detecting potential errors in human rating, or surfacing critical context to support human rating. We outline how to support all of these design patterns using a single, optimized prompt. Beyond these synthetic experiments, we share how piloting our proposed techniques in a real-world review queue yielded a 41.5% improvement in optimizing available human rater capacity, and a 9–11 % increase (absolute) in precision and recall for detecting violative content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations CentresRonal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mohan Baruwal Chhetri et al.S&P 2026 · 44 citations
- How Generative AI Empowers Attackers and Defenders Across the Trust & Safety LandscapePatrick Gage Kelley, Steven Rousso-Schindler, Renee Shelby, Kurt Thomas et al.CHI 2026 · 4 citations
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl et al.CHI 2026 · 3 citations
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- Test-Time Detoxification without Training or Learning AnythingBaturay Saglam, Dionysios KalogeriasICML 2026 · 2 citations
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- SoK: Hate, Harassment, and the Changing Landscape of Online AbuseKurt Thomas, Devdatta Akhawe, Michael D. Bailey, Dan Boneh et al.S&P 2021 · 175 citations
Related papers
- Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media PlatformsRajvardhan Oak, Muhammad Haroon, Claire Wonjeong Jo, Magdalena Wojcieszak et al.ACL 2025 · 1 citation
- Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less UsefulChenyan Jia, Apoorva Gondimalla, Angie Zhang, David Joseph Mullings et al.CHI 2026 · 1 citation
- Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé et al.EMNLP 2025
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek et al.EMNLP 2024 · 4 citations
- Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsZhuoran Lu, Gionnieve Lim, Ming YinCHI 2026
