MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
Han Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei Lee
Abstract
Hate speech is a pressing issue in modern society, with significant effects both online and offline. Recent research in hate speech detection has primarily centered on text-based media, largely overlooking multimodal content such as videos. Existing studies on hateful video datasets have predominantly focused on English content within a Western context and have been limited to binary labels (hateful or non-hateful), lacking detailed contextual information. This study presents MultiHateClip 1 , an novel multilingual dataset created through hate lexicons and human annotation. It aims to enhance the detection of hateful videos on platforms such as YouTube and Bilibili, including content in both English and Chinese languages. Comprising 2,000 videos annotated for hatefulness, offensiveness, and normalcy, this dataset provides a cross-cultural perspective on gender-based hate speech. Through a detailed examination of human annotation results, we discuss the differences between Chinese and English hateful videos and underscore the importance of different modalities in hateful and offensive video analysis. Evaluations of state-of-the-art video classification models, such as VLM, GPT-4V and Qwen-VL, on MultiHateClip highlight the existing challenges in accurately distinguishing between hateful and offensive content and the urgent need for models that are both multimodally and culturally nuanced. MultiHateClip represents a foundational advance in enhancing hateful video detection by underscoring the necessity of a multimodal and culturally sensitive approach in combating online hate speech.
Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate DetectionJian Lang, Rongpei Hong, Jin Xu, Yili Li et al.WWW 2025 · 14 citations
- ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in VideosMohammad Zia Ur Rehman, Anukriti Bhatnagar, Omkar Kabde, Shubhi Bansal et al.ACL 2025 · 11 citations
- MM-HSD: Multi-Modal Hate Speech Detection in VideosBerta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko et al.ACM MM 2025 · 5 citations
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang et al.WWW 2026 · 3 citations
- Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection Via Cross-Domain Retrieval AugmentationRongpei Hong, Jian Lang, Ting Zhong, Fan ZhouICCV 2025 · 3 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Understanding and Evaluating Racial Biases in Image CaptioningDora Zhao, Angelina Wang, Olga RussakovskyICCV 2021 · 165 citations
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng et al.EMNLP 2022 · 82 citations
Related papers
- MemeCLIP: Leveraging CLIP Representations for Multimodal Meme ClassificationSiddhant Bikram Shah, Shuvam Shiwakoti, Maheep Chaudhary, Haohan WangEMNLP 2024 · 12 citations
- HVGuard: Utilizing Multimodal Large Language Models for Hateful Video DetectionYiheng Jing, Mingming Zhang, Yong Zhuang, Jiacheng Guo et al.EMNLP 2025 · 1 citation
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 11 citations
- ChinaOpen: A Dataset for Open-world Multimodal LearningAozhu Chen, Ziyuan Wang, Chengbo Dong, Kaibin Tian et al.ACM MM 2023 · 8 citations
- Disentangling Hate in Online MemesRoy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang et al.ACM MM 2021 · 85 citations
