Mitigating Bias in Session-based Cyberbullying Detection: A Non-Compromising Approach
Lu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall, Huan Liu
Abstract
The element of repetition in cyberbullying behavior has directed recent computational studies toward detecting cyberbullying based on a social media session. In contrast to a single text, a session may consist of an initial post and an associated sequence of comments. Yet, emerging efforts to enhance the performance of session-based cyberbullying detection have largely overlooked unintended social biases in existing cyberbullying datasets. For example, a session containing certain demographicidentity terms (e.g., "gay" or "black") is more likely to be classified as an instance of cyberbullying. In this paper, we first show evidence of such bias in models trained on sessions collected from different social media platforms (e.g., Instagram). We then propose a contextaware and model-agnostic debiasing strategy that leverages a reinforcement learning technique, without requiring any extra resources or annotations apart from a pre-defined set of sensitive triggers commonly used for identifying cyberbullying instances. Empirical evaluations show that the proposed strategy can simultaneously alleviate the impacts of the unintended biases and improve the detection performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 521f32d2-c6f6-4aeb-b53a-909cd873ed01Cited by top-tier papers4
- A Human-Centered Systematic Literature Review of the Computational Approaches for Online Sexual Risk DetectionAfsaneh Razi, Seunghyun Kim, Ashwaq Alsoubai, Gianluca Stringhini et al.CSCW 2021 · 93 citations
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 11 citations
- Bias Mitigation for Toxicity Detection via Sequential DecisionsLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall et al.SIGIR 2022 · 8 citations
- CoNewsReader: Supporting Comprehensive Understanding and Raising Critical Thoughts on Social Media News Through CommentsKangyu Yuan, Guanzheng Chen, Sizhe Liang, Hehai Lin et al.CSCW 2026
Builds on2
- A Reinforced Generation of Adversarial Examples for Neural Machine TranslationWei Zou, Shujian Huang, Jun Xie, Xinyu Dai et al.ACL 2020 · 66 citations
- Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance WeightingGuanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai et al.ACL 2020 · 56 citations
Related papers
- Improving Cyberbullying Detection with User InteractionSuyu Ge, Lu Cheng, Huan LiuWWW 2021 · 41 citations
- HENIN: Learning Heterogeneous Neural Interaction Networks for Explainable Cyberbullying Detection on Social MediaHsin-Yu Chen, Cheng-Te LiEMNLP 2020 · 27 citations
- Automated Detection of Doxing on TwitterYounes Karimi, Anna Cinzia Squicciarini, Shomir WilsonCSCW 2022 · 18 citations
- GenEx: A Commonsense-aware Unified Generative Framework for Explainable Cyberbullying DetectionKrishanu Maity, Raghav Jain, Prince Jha, Sriparna Saha et al.EMNLP 2023 · 4 citations
- Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network AbuseGarrett Wilson, Geoffrey Goh, Yan Jiang, Ajay Gupta et al.USENIX Security 2025
