Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection
Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, Fan Zhou
摘要
Short Video Hate Detection (SVHD) is increasingly vital as hateful content - such as racial and gender-based discrimination - spreads rapidly across platforms like TikTok, YouTube Shorts, and Instagram Reels. Existing approaches face significant challenges: hate expressions continuously evolve, hateful signals are dispersed across multiple modalities (audio, text, and vision), and the contribution of each modality varies across different hate content. To address these issues, we introduce MoRE(Mixture of Retrieval-augmented multimodal Experts), a novel framework designed to enhance SVHD. MoRE employs specialized multimodal experts for each modality, leveraging their unique strengths to identify hateful content effectively. To ensure model's adaptability to rapidly evolving hate content, MoRE leverages contextual knowledge extracted from relevant instances retrieved by a powerful joint multimodal video retriever for each target short video. Moreover, a dynamic sample-sensitive integration network adaptively adjusts the importance of each modality on a per-sample basis, optimizing the detection process by prioritizing the most informative modalities for each instance. Our MoRE adopts an end-to-end training strategy that jointly optimizes both expert networks and the overall framework, resulting in nearly a twofold improvement in training efficiency, which in turn enhances its applicability to real-world scenarios. Extensive experiments on three benchmarks demonstrate that MoRE surpasses state-of-the-art baselines, achieving an average improvement of 6.91% in macro-F1 score across all datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MM-HSD: Multi-Modal Hate Speech Detection in VideosBerta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko 等ACM MM 2025 · 被引用 5 次
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang 等WWW 2026 · 被引用 3 次
- Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection Via Cross-Domain Retrieval AugmentationRongpei Hong, Jian Lang, Ting Zhong, Fan ZhouICCV 2025 · 被引用 3 次
- Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time AdaptationJiao Li, Jian Lang, Xikai Tang, Wenzheng Shu 等AAAI 2026
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li 等ACL 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
相关 Paper
- MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking PerturbationsQiyao Xue, Yuchen Dou, Zheyuan Ryan Shi, Xiang Lorraine Li 等AAAI 2026 · 被引用 2 次
- Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful VideosJunyu Lu, Deyi Ji, Liqun Liu, Xiaokun Zhang 等SIGIR 2026
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu 等ACM MM 2024 · 被引用 43 次
- Detecting Fake News in Short Videos Through Multi-View AggregationNuo Li, Yuan Xiong, Chengliang Liu, Jie Wen 等AAAI 2026
- HVGuard: Utilizing Multimodal Large Language Models for Hateful Video DetectionYiheng Jing, Mingming Zhang, Yong Zhuang, Jiacheng Guo 等EMNLP 2025 · 被引用 1 次
