ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
Mohammad Zia Ur Rehman, Anukriti Bhatnagar, Omkar Kabde, Shubhi Bansal, Nagendra Kumar
摘要
The existing research has primarily focused on text and image-based hate speech detection, video-based approaches remain underexplored. In this work, we introduce a novel dataset, Im-pliHateVid, specifically curated for implicit hate speech detection in videos. ImpliHate-Vid consists of 2,009 videos comprising 509 implicit hate videos, 500 explicit hate videos, and 1,000 non-hate videos, making it one of the first large-scale video datasets dedicated to implicit hate detection. We also propose a novel two-stage contrastive learning framework for hate speech detection in videos. In the first stage, we train modality-specific encoders for audio, text, and image using contrastive loss by concatenating features from the three encoders. In the second stage, we train cross-encoders using contrastive learning to refine multimodal representations. Additionally, we incorporate sentiment, emotion, and caption-based features to enhance implicit hate detection. We evaluate our method on two datasets, ImpliHate-Vid for implicit hate speech detection and another dataset for general hate speech detection in videos, HateMM dataset, demonstrating the effectiveness of the proposed multimodal contrastive learning for hateful content detection in videos and the significance of our dataset. The code and dataset will be made available on the GitHub repository 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li 等ACL 2026
- Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful VideosJunyu Lu, Deyi Ji, Liqun Liu, Xiaokun Zhang 等SIGIR 2026
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
相关 Paper
- MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and BilibiliHan Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei LeeACM MM 2024 · 被引用 23 次
- MM-HSD: Multi-Modal Hate Speech Detection in VideosBerta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko 等ACM MM 2025 · 被引用 5 次
- HVGuard: Utilizing Multimodal Large Language Models for Hateful Video DetectionYiheng Jing, Mingming Zhang, Yong Zhuang, Jiacheng Guo 等EMNLP 2025 · 被引用 1 次
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate DetectionYejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub HanEMNLP 2025 · 被引用 2 次
- HateProof: Are Hateful Meme Detection Systems really Robust?Piush Aggarwal, Pranit Chawla, Mithun Das, Punyajoy Saha 等WWW 2023 · 被引用 15 次
