Hate-Speech and Offensive Language Detection in Roman Urdu
Hammad Rizwan, Muhammad Haroon Shakeel, Asim Karim
摘要
The task of automatic hate-speech and offensive language detection in social media content is of utmost importance due to its implications in unprejudiced society concerning race, gender, or religion. Existing research in this area, however, is mainly focused on the English language, limiting the applicability to particular demographics. Despite its prevalence, Roman Urdu (RU) lacks language resources, annotated datasets, and language models for this task. In this study, we: (1) Present a lexicon of hateful words in RU, (2) Develop an annotated dataset called RUHSOLD consisting of 10, 012 tweets in RU with both coarse-grained and fine-grained labels of hate-speech and offensive language, (3) Explore the feasibility of transfer learning of five existing embedding models to RU, (4) Propose a novel deep learning architecture called CNN-gram for hatespeech and offensive language detection and compare its performance with seven current baseline approaches on RUHSOLD dataset, and (5) Train domain-specific embeddings on more than 4.7 million tweets and make them publicly available. We conclude that transfer learning is more beneficial as compared to training embedding from scratch and that the proposed model exhibits greater robustness as compared to the baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn 等EMNLP 2022 · 被引用 41 次
- Detecting Propaganda Techniques in Code-Switched Social Media TextMuhammad Umar Salman, Asif Hanif, Shady Shehata, Preslav NakovEMNLP 2023 · 被引用 5 次
- BeyondGender: A Multifaceted Bilingual Dataset for Practical Sexism DetectionXuan Luo, Li Yang, Han Zhang, Geng Tu 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 被引用 11 次
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq 等ACL 2024
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini 等EMNLP 2021 · 被引用 2 次
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis 等ACL 2021
