Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection
Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, Fan Zhou
Abstract
Short Video Hate Detection (SVHD) is increasingly vital as hateful content - such as racial and gender-based discrimination - spreads rapidly across platforms like TikTok, YouTube Shorts, and Instagram Reels. Existing approaches face significant challenges: hate expressions continuously evolve, hateful signals are dispersed across multiple modalities (audio, text, and vision), and the contribution of each modality varies across different hate content. To address these issues, we introduce MoRE(Mixture of Retrieval-augmented multimodal Experts), a novel framework designed to enhance SVHD. MoRE employs specialized multimodal experts for each modality, leveraging their unique strengths to identify hateful content effectively. To ensure model's adaptability to rapidly evolving hate content, MoRE leverages contextual knowledge extracted from relevant instances retrieved by a powerful joint multimodal video retriever for each target short video. Moreover, a dynamic sample-sensitive integration network adaptively adjusts the importance of each modality on a per-sample basis, optimizing the detection process by prioritizing the most informative modalities for each instance. Our MoRE adopts an end-to-end training strategy that jointly optimizes both expert networks and the overall framework, resulting in nearly a twofold improvement in training efficiency, which in turn enhances its applicability to real-world scenarios. Extensive experiments on three benchmarks demonstrate that MoRE surpasses state-of-the-art baselines, achieving an average improvement of 6.91% in macro-F1 score across all datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb1a1b00-bf71-4c0f-a1a3-14227f0ef0cfCited by top-tier papers9
- MM-HSD: Multi-Modal Hate Speech Detection in VideosBerta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko et al.ACM MM 2025 · 5 citations
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang et al.WWW 2026 · 3 citations
- Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection Via Cross-Domain Retrieval AugmentationRongpei Hong, Jian Lang, Ting Zhong, Fan ZhouICCV 2025 · 3 citations
- Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time AdaptationJiao Li, Jian Lang, Xikai Tang, Wenzheng Shu et al.AAAI 2026
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li et al.ACL 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking PerturbationsQiyao Xue, Yuchen Dou, Zheyuan Ryan Shi, Xiang Lorraine Li et al.AAAI 2026 · 2 citations
- Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful VideosJunyu Lu, Deyi Ji, Liqun Liu, Xiaokun Zhang et al.SIGIR 2026
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu et al.ACM MM 2024 · 43 citations
- Detecting Fake News in Short Videos Through Multi-View AggregationNuo Li, Yuan Xiong, Chengliang Liu, Jie Wen et al.AAAI 2026
- HVGuard: Utilizing Multimodal Large Language Models for Hateful Video DetectionYiheng Jing, Mingming Zhang, Yong Zhuang, Jiacheng Guo et al.EMNLP 2025 · 1 citation
