Wait! There's a Way Out: A Decision Mechanism for Forecasting Conversational Derailment
Laerdon Kim, Vivian Nguyen, Cristian Danescu-Niculescu-Mizil
Abstract
Forecasting conversational derailment is the task of predicting, as the conversation unfolds, whether it will eventually derail into personal attacks. Since forecasting models operate in an online fashion, they must decide whether to "trigger" an alert after each utterance-for example, to notify participants or a moderator that the conversation is at risk of derailing. Existing approaches make this decision solely based on the estimated likelihood of derailment given the preceding utterances, implicitly assuming that the conversation's future trajectory is fixed. As a result, they ignore the possibility of future recovery and incur an unnecessarily high rate of false positives. In this work we propose a method for decoupling the decision to trigger from the derailment likelihood estimation. Our approach is inspired by the first human baseline on this task, which shows that humans achieve dramatically lower false positive rates by selectively deferring their decision to trigger when they anticipate that tension is likely to subside. We operationalize this insight with a deferral mechanism that uses forward-looking simulations to assess whether a tense moment admits plausible paths to recovery. Incorporating this mechanism into a state-of-the-art forecasting model substantially reduces false positives without sacrificing forecasting accuracy. More broadly, this work highlights the value of treating decision making as a first-class component of forecasting systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7c17b0c-4837-4fb1-80aa-518b4c593ae2Builds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Conversations Gone Alright: Quantifying and Predicting Prosocial Outcomes in Online ConversationsJiajun Bao, Junjie Wu, Yiming Zhang, Eshwar Chandrasekharan et al.WWW 2021 · 63 citations
- Proactive Moderation of Online Discussions: Existing Practices and the Potential for Algorithmic SupportCharlotte Schluger, Jonathan P. Chang, Cristian Danescu-Niculescu-Mizil, Karen LevyCSCW 2022 · 41 citations
- Thread With Caution: Proactively Helping Users Assess and Deescalate Tension in Their Online DiscussionsJonathan P. Chang, Charlotte Schluger, Cristian Danescu-Niculescu-MizilCSCW 2022 · 27 citations
- Hanging in the Balance: Pivotal Moments in Crisis Counseling ConversationsVivian Nguyen, Lillian Lee, Cristian Danescu-Niculescu-MizilACL 2025 · 2 citations
Related papers
- A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online ConversationsYingxue Fu, Anaïs OllagnierACL 2026
- Toxicity Ahead: Forecasting Conversational Derailment on GitHubMia Mohammad Imran, Robert Zita, Rahat Rizvi Rahman, Preetha Chatterjee et al.ICSE 2026
- Differentiable Learning Under TriageNastaran Okati, Abir De, Manuel Gomez-RodriguezNeurIPS 2021 · 99 citations
- Don't Let Me Be Misunderstood: Comparing Intentions and Perceptions in Online DiscussionsJonathan P. Chang, Justin Cheng, Cristian Danescu-Niculescu-MizilWWW 2020 · 29 citations
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 62 citations
