Supporting Human Raters with the Detection of Harmful Content Using Large Language Models
Kurt Thomas, Patrick Gage Kelley, David Tao, Sarah Meiklejohn, Owen Vallis, Shunwen Tan, Blaz Bratanic, Felipe Tiengo Ferreira, Vijay Kumar Eranti, Elie Bursztein
摘要
In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election misinformation. Using a dataset of 50,000 user comments, we demonstrate that LLMs can achieve 90 % accuracy when compared to human verdicts. We explore how to best leverage these capabilities, proposing five design patterns that integrate LLMs with human rating, such as pre-filtering non-violative content, detecting potential errors in human rating, or surfacing critical context to support human rating. We outline how to support all of these design patterns using a single, optimized prompt. Beyond these synthetic experiments, we share how piloting our proposed techniques in a real-world review queue yielded a 41.5% improvement in optimizing available human rater capacity, and a 9–11 % increase (absolute) in precision and recall for detecting violative content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations CentresRonal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mohan Baruwal Chhetri 等S&P 2026 · 被引用 44 次
- How Generative AI Empowers Attackers and Defenders Across the Trust & Safety LandscapePatrick Gage Kelley, Steven Rousso-Schindler, Renee Shelby, Kurt Thomas 等CHI 2026 · 被引用 4 次
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl 等CHI 2026 · 被引用 3 次
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 被引用 2 次
- Test-Time Detoxification without Training or Learning AnythingBaturay Saglam, Dionysios KalogeriasICML 2026 · 被引用 2 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster 等ICLR 2023 · 被引用 297 次
- SoK: Hate, Harassment, and the Changing Landscape of Online AbuseKurt Thomas, Devdatta Akhawe, Michael D. Bailey, Dan Boneh 等S&P 2021 · 被引用 175 次
相关 Paper
- Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media PlatformsRajvardhan Oak, Muhammad Haroon, Claire Wonjeong Jo, Magdalena Wojcieszak 等ACL 2025 · 被引用 1 次
- Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less UsefulChenyan Jia, Apoorva Gondimalla, Angie Zhang, David Joseph Mullings 等CHI 2026 · 被引用 1 次
- Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé 等EMNLP 2025
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
- Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsZhuoran Lu, Gionnieve Lim, Ming YinCHI 2026
