ToxiShield: Promoting Inclusive Developer Communication through Real-Time Toxicity Filtering
Md Awsaf Alam Anindya, Showvik Biswas, Anindya Iqbal, Jaydeb Sarker, Amiangshu Bosu
Abstract
Toxic interactions during code reviews can undermine teamwork and hinder productivity in software engineering (SE) teams. While prior studies explore toxicity detection and empirical investigation, they lack real-time detoxification tools to support the SE community. To address this gap, we present ToxiShield, a browser extension for GitHub pull requests that is built using three modules: i) Toxicity Filter -to identify whether a text is toxic, ii) Communication coach -to facilitate just-in-time fine-grained toxicity categorization with explanations, and iii) The Reframer -that generates a revised, constructive alternative of a toxic text. For each module, we trained and evaluated multiple deep learning and Large Language Models (LLMs) to identify the best choice. A BERT-based binary detection model, trained on 38,761 code review samples, achieves 98% accuracy and an F1-score of 97% and is the selected one for the Toxicity Filter module. For the Communication Coach, prompt-tuned Claude 3.5 Sonnet achieved the best performance with 39%𝑀𝐶𝐶 and 42%𝐹 1 in multiclass toxicity classification with detailed reasoning. For Reframer, we evaluated five LLMs using a fine-tuning strategy on a dataset of 10,120 code review comments. The fine-tuned Llama 3.2 model achieves 95.27% style transfer accuracy, 97.03% fluency, 67.07% content preservation, and an 84% J-score. We further validated ToxiShield through a human evaluation using the Technology Acceptance Model with 10 participants, confirming its perceived usefulness and ease of adoption. ToxiShield sets a benchmark for advancing constructive communication in software engineering, driving inclusivity and healthier collaboration in open-source communities.
CCS Concepts: • Software and its engineering → Collaboration in software development; • Collaboration in software development → Detoxification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 058b3ae4-1dab-4413-a020-c56d37d09edbBuilds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- "Did You Miss My Comment or What?" Understanding Toxicity in Open Source DiscussionsCourtney Miller, Sophie Cohen, Daniel Klug, Bogdan Vasilescu et al.ICSE 2022 · 54 citations
- The "Shut the f**k up" Phenomenon: Characterizing Incivility in Open Source Code Review DiscussionsIsabella Ferreira, Jinghui Cheng, Bram AdamsCSCW 2021 · 51 citations
Related papers
- Toxicity Ahead: Forecasting Conversational Derailment on GitHubMia Mohammad Imran, Robert Zita, Rahat Rizvi Rahman, Preetha Chatterjee et al.ICSE 2026
- RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive VisualizationAustin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson et al.CSCW 2021 · 32 citations
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang et al.EMNLP 2023 · 12 citations
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng et al.ACL 2026
