RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive Visualization
Austin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson, Muhammed Ahmed, Stephane Pinel, Duen Horng (Polo) Chau, Diyi Yang
摘要
With the widespread use of toxic language online, platforms are increasingly using automated systems that leverage advances in natural language processing to automatically flag and remove toxic comments. However, most automated systems-when detecting and moderating toxic language-do not provide feedback to their users, let alone provide an avenue of recourse for these users to make actionable changes. We present our work, Recast, an interactive, open-sourced web tool for visualizing these models' toxic predictions, while providing alternative suggestions for flagged toxic language. Our work also provides users with a new path of recourse when using these automated moderation tools. Recast highlights text responsible for classifying toxicity, and allows users to interactively substitute potentially toxic phrases with neutral alternatives. We examined the effect of Recast via two large-scale user evaluations, and found that Recast was highly effective at helping users reduce toxicity as detected through the model. Users also gained a stronger understanding of the underlying toxicity criterion used by black-box models, enabling transparency and recourse. In addition, we found that when users focus on optimizing language for these models instead of their own judgement (which is the implied incentive and goal of deploying automated models), these models cease to be effective classifiers of toxicity compared to human annotations. This opens a discussion for how toxicity detection models work and should work, and their effect on the future of online discourse.
CCS Concepts: • Human-centered computing → Collaborative and social computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- End-User Audits: A System Empowering Communities to Lead Large-Scale Investigations of Harmful Algorithmic BehaviorMichelle S. Lam, Mitchell L. Gordon, Danaë Metaxa, Jeffrey T. Hancock 等CSCW 2022 · 被引用 77 次
- Farsight: Fostering Responsible AI Awareness During AI Application PrototypingZijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry 等CHI 2024 · 被引用 55 次
- Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision MakingSara Salimzadeh, Gaole He, Ujwal GadirajuCHI 2024 · 被引用 40 次
- "I Got Flagged for Supposed Bullying, Even Though It Was in Response to Someone Harassing Me About My Disability.": A Study of Blind TikTokers' Content Moderation ExperiencesYao Lyu, Jie Cai, Anisa Callis, Kelley Cotter 等CHI 2024 · 被引用 25 次
- Hate Raids on Twitch: Understanding Real-Time Human-Bot Coordinated Attacks in Live Streaming CommunitiesJie Cai, Sagnik Chowdhury, Hongyang Zhou, Donghee Yvette WohnCSCW 2023 · 被引用 17 次
它引用的顶会 Paper3
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- Reconsidering Self-Moderation: the Role of Research in Supporting Community-Based Models for Online Content ModerationJoseph SeeringCSCW 2020 · 被引用 151 次
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based SystemsC. Estelle Smith, Bowen Yu, Anjali Srivastava, Aaron Halfaker 等CHI 2020 · 被引用 77 次
相关 Paper
- ToxiShield: Promoting Inclusive Developer Communication through Real-Time Toxicity FilteringMd Awsaf Alam Anindya, Showvik Biswas, Anindya Iqbal, Jaydeb Sarker 等FSE 2026 · 被引用 1 次
- Thread With Caution: Proactively Helping Users Assess and Deescalate Tension in Their Online DiscussionsJonathan P. Chang, Charlotte Schluger, Cristian Danescu-Niculescu-MizilCSCW 2022 · 被引用 27 次
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 被引用 74 次
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi 等EMNLP 2025
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 被引用 7 次
