Toxicity Ahead: Forecasting Conversational Derailment on GitHub
Mia Mohammad Imran, Robert Zita, Rahat Rizvi Rahman, Preetha Chatterjee, Kostadin Damevski
Abstract
Toxic interactions in Open Source Software (OSS) communities reduce contributor engagement and threaten project sustainability. Preventing such toxicity before it emerges requires a clear understanding of how harmful conversations unfold. However, most proactive moderation strategies are manual, requiring significant time and effort from community maintainers. To support more scalable approaches, we curate a dataset of 159 derailed toxic threads and 207 non-toxic threads from GitHub discussions. Our analysis reveals that toxicity can be forecasted by tension triggers, sentiment shifts, and specific conversational patterns.
We present a novel Large Language Model (LLM)-based framework for predicting conversational derailment on GitHub using a two-step prompting pipeline. First, we generate Summaries of Conversation Dynamics (SCDs) via Least-to-Most (LtM) prompting; then we use these summaries to estimate the likelihood of derailment. Evaluated on Qwen and Llama models, our LtM strategy achieves F1-scores of 0.901 and 0.852, respectively, at a decision threshold of 0.3, outperforming established NLP baselines on conversation derailment. External validation on a dataset of 308 GitHub issue threads (65 toxic, 243 non-toxic) yields an F1-score up to 0.797. Our findings demonstrate the effectiveness of structured LLM prompting for early detection of conversational derailment in OSS, enabling proactive and explainable moderation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3888848-3739-4969-8db6-83b69bee5960Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu et al.ICLR 2023 · 94 citations
- Conversations Gone Alright: Quantifying and Predicting Prosocial Outcomes in Online ConversationsJiajun Bao, Junjie Wu, Yiming Zhang, Eshwar Chandrasekharan et al.WWW 2021 · 63 citations
- Code of Conduct Conversations in Open Source Software Projects on GithubRenee Li, Pavitthra Pandurangan, Hana Frluckaj, Laura DabbishCSCW 2021 · 49 citations
Related papers
- A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online ConversationsYingxue Fu, Anaïs OllagnierACL 2026
- Toxicity Detection for FreeZhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao et al.NeurIPS 2024 · 20 citations
- Beyond Adoption: Examining the Evolution and Impact of Code of Conduct on Open Source CommunitiesJiayi Sun, Hongbo Fang, Junming Zhang, Jiakai Shi et al.ICSE 2026 · 2 citations
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi Xie, Minghong Fang, Renjie Pi, Neil GongACL 2024
- The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHubJaydeb Sarker, Asif Kamal Turzo, Amiangshu BosuFSE 2025 · 1 citation
