How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
Patrick Gage Kelley, Steven Rousso-Schindler, Renee Shelby, Kurt Thomas, Allison Woodruff
Abstract
Generative AI (GenAI) is a powerful technology poised to reshape Trust & Safety. While misuse by attackers is a growing concern, its defensive capacity remains underexplored. This paper examines these effects through a qualitative study with 43 Trust & Safety experts across five domains: child safety, election integrity, hate and harassment, scams, and violent extremism. Our findings characterize a landscape in which GenAI empowers both attackers and defenders. GenAI dramatically increases the scale and speed of attacks, lowering the barrier to entry for creating harmful content, including sophisticated propaganda and deepfakes. Conversely, defenders envision leveraging GenAI to detect and mitigate harmful content at scale, conduct investigations, deploy persuasive counternarratives, improve moderator wellbeing, and offer user support. This work provides a strategic framework for understanding GenAI's impact on Trust & Safety and charts a path for its responsible use in creating safer online environments.
• Human-centered computing → Human computer interaction (HCI); • Social and professional topics → Computing / technology policy; • Security and privacy → Human and societal aspects of security and privacy; • Computing methodologies → Artificial intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b69c8192-1746-44db-9e1e-292201400834Builds on26
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for Shifting Organizational PracticesBogdana Rakova, Jingying Yang, Henriette Cramer, Rumman ChowdhuryCSCW 2021 · 326 citations
- SoK: Hate, Harassment, and the Changing Landscape of Online AbuseKurt Thomas, Devdatta Akhawe, Michael D. Bailey, Dan Boneh et al.S&P 2021 · 175 citations
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportMichael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan et al.CSCW 2022 · 149 citations
Related papers
- Who Gets to Define Safety? A Systematic Review of How Generative AI Research Addresses Youth Online SafetyOzioma Collins Oguine, Adriana Alvarado Garcia, Michael J. Muller, Karla Badillo-UrquiolaCHI 2026 · 2 citations
- How Tech Workers Contend with Hazards of Humanlikeness in Generative AIMark Diaz, Renee Shelby, Eric Corbett, Andrew SmartCHI 2026 · 1 citation
- Families' Vision of Generative AI Agents for Household Safety Against Digital and Physical ThreatsZikai Wen, Lanjing Liu, Yaxing YaoCSCW 2025 · 3 citations
- A Framework to Characterize Reporting on Generative AI UseAgathe Balayn, Varun Nagaraj Rao, Su Lin Blodgett, Aylin Caliskan et al.CHI 2026 · 1 citation
- From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating DisordersAmy A. Winecoff, Kevin KlymanCHI 2026
