ACL2026

BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation

Faisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md. Tahmid Hasan Fuad, Md Farhan Ishmam, Faria Sultana, M. Ashraful Amin, Amin Ahsan Ali, Akmmahbubur Rahman

摘要

Online safety in low-resource languages hinges not only on accurate hate speech detection but also on transparent, culturally grounded explanations. Yet prior work in Bangla largely focuses on hate classification while overlooking interpretability. We address this gap by introducing BANHADEX, the first hate explainability dataset in Bangla with human-annotated labels. BANHADEX contains 19,203 YouTube comments spanning April 2024-June 2025, annotated for binary hate classification with seven fine-grained hate categories, seven target groups, and concise explanations for each sample. Our data pipeline relies on a two-stage annotation protocol that uses majority voting for robust labeling. Our rich suite of experiments on open and closed-source LLMs reveals that explanation-guided LoRA substantially outperforms both classification and explanation quality across prompting and finetuning strategies. BANHADEX establishes the groundwork for faithful interpretability and safer moderation in linguistically rich yet underresourced languages. The code and dataset are publicly available at: https://github.com/ MOSHIIUR/BANHADEX .