RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive Visualization
Austin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson, Muhammed Ahmed, Stephane Pinel, Duen Horng (Polo) Chau, Diyi Yang
Abstract
With the widespread use of toxic language online, platforms are increasingly using automated systems that leverage advances in natural language processing to automatically flag and remove toxic comments. However, most automated systems-when detecting and moderating toxic language-do not provide feedback to their users, let alone provide an avenue of recourse for these users to make actionable changes. We present our work, Recast, an interactive, open-sourced web tool for visualizing these models' toxic predictions, while providing alternative suggestions for flagged toxic language. Our work also provides users with a new path of recourse when using these automated moderation tools. Recast highlights text responsible for classifying toxicity, and allows users to interactively substitute potentially toxic phrases with neutral alternatives. We examined the effect of Recast via two large-scale user evaluations, and found that Recast was highly effective at helping users reduce toxicity as detected through the model. Users also gained a stronger understanding of the underlying toxicity criterion used by black-box models, enabling transparency and recourse. In addition, we found that when users focus on optimizing language for these models instead of their own judgement (which is the implied incentive and goal of deploying automated models), these models cease to be effective classifiers of toxicity compared to human annotations. This opens a discussion for how toxicity detection models work and should work, and their effect on the future of online discourse.
CCS Concepts: • Human-centered computing → Collaborative and social computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b4b3fd7-7c86-4e89-b2c1-0cb080c03270Cited by top-tier papers12
- End-User Audits: A System Empowering Communities to Lead Large-Scale Investigations of Harmful Algorithmic BehaviorMichelle S. Lam, Mitchell L. Gordon, Danaë Metaxa, Jeffrey T. Hancock et al.CSCW 2022 · 77 citations
- Farsight: Fostering Responsible AI Awareness During AI Application PrototypingZijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry et al.CHI 2024 · 55 citations
- Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision MakingSara Salimzadeh, Gaole He, Ujwal GadirajuCHI 2024 · 40 citations
- "I Got Flagged for Supposed Bullying, Even Though It Was in Response to Someone Harassing Me About My Disability.": A Study of Blind TikTokers' Content Moderation ExperiencesYao Lyu, Jie Cai, Anisa Callis, Kelley Cotter et al.CHI 2024 · 25 citations
- Hate Raids on Twitch: Understanding Real-Time Human-Bot Coordinated Attacks in Live Streaming CommunitiesJie Cai, Sagnik Chowdhury, Hongyang Zhou, Donghee Yvette WohnCSCW 2023 · 17 citations
Builds on3
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Reconsidering Self-Moderation: the Role of Research in Supporting Community-Based Models for Online Content ModerationJoseph SeeringCSCW 2020 · 151 citations
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based SystemsC. Estelle Smith, Bowen Yu, Anjali Srivastava, Aaron Halfaker et al.CHI 2020 · 77 citations
Related papers
- ToxiShield: Promoting Inclusive Developer Communication through Real-Time Toxicity FilteringMd Awsaf Alam Anindya, Showvik Biswas, Anindya Iqbal, Jaydeb Sarker et al.FSE 2026 · 1 citation
- Thread With Caution: Proactively Helping Users Assess and Deescalate Tension in Their Online DiscussionsJonathan P. Chang, Charlotte Schluger, Cristian Danescu-Niculescu-MizilCSCW 2022 · 27 citations
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 74 citations
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi et al.EMNLP 2025
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
