LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
Md. Arid Hasan, Firoj Alam, Md Fahad Hossain, Usman Naseem, Syed Ishtiaque Ahmed
Abstract
Online social media platforms are central to everyday communication and information seeking. While these platforms serve positive purposes, they also provide fertile ground for the spread of hate speech, offensive language, and bullying content targeting individuals, organizations, and communities. Such content undermines safety, participation, and equity online. Reliable detection systems are therefore needed, especially for low-resource languages where moderation tools are limited. In Bangla, prior work has contributed resources and models, but most are single-task (e.g., binary hate/offense) with limited coverage of multi-facet signals (type, severity, target). We address these gaps by introducing the first multitask Bangla hate-speech dataset, BanglaMul-tiHate, one of the largest manually annotated corpus to date. Building on this resource, we conduct a comprehensive, controlled comparison spanning classical baselines, monolingual pretrained models, and LLMs under zero-shot prompting and LoRA fine-tuning. Our experiments assess LLM adaptability in a lowresource setting and reveal a consistent trend: although LoRA-tuned LLMs are competitive with BanglaBERT, culturally and linguistically grounded pretraining remains critical for robust performance. Together, our dataset and findings establish a stronger benchmark for developing culturally aligned moderation tools in low-resource contexts. For reproducibility, we will release the dataset and all related scripts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0350453e-932c-41b3-8673-266e19b0fa5cBuilds on2
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
Related papers
- BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanationFaisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md. Tahmid Hasan Fuad et al.ACL 2026
- BANMIME : Misogyny Detection with Metaphor Explanation on Bangla MemesMd Ayon Mia, Akm Moshiur Rahman Mazumder, Khadiza Sultana Sayma, Md Fahim et al.EMNLP 2025
- Deciphering Hate: Identifying Hateful Memes and Their TargetsEftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah Masud PreumACL 2024 · 9 citations
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq et al.ACL 2024
- BanglaAbuseMeme: A Dataset for Bengali Abusive Meme ClassificationMithun Das, Animesh MukherjeeEMNLP 2023 · 9 citations
