Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection Via Cross-Domain Retrieval Augmentation
Rongpei Hong, Jian Lang, Ting Zhong, Fan Zhou
摘要
The rapid proliferation of online video-sharing platforms has accelerated the spread of malicious videos, creating an urgent need for robust detection methods. However, the performance and generalizability of existing detection approaches are severely limited by the scarcity of annotated video data, as manually curating large-scale malicious detection datasets is both labor-intensive and impractical. To address this challenge, we propose CRAVE, a novel CRoss-domAin retrieVal augmEntation framework that transfers knowledge from resource-rich image-text domain to enhance malicious video detection. Specifically, CRAVE introduces a Pseudo-Pair Retriever to identify semantically relevant image-text data for high-quality cross-domain augmentation. Additionally, a Contrastive Cross-Domain Augmenter is designed to disentangle domain-shared and -unique representations, effectively bridging the domain gaps during knowledge transfer. These shared image-text representations are then leveraged to refine video representations, yielding more discriminative features for accurate malicious content detection. Experiments on four video datasets demonstrate that CRAVE largely outperforms competitive baselines in both performance and generalization, providing an innovative and strong solution to the issue of video data-scarcity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt TuningJian Lang, Hong, Ting Zhong, Fan ZhouICML 2026 · 被引用 1 次
- Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time AdaptationJiao Li, Jian Lang, Xikai Tang, Wenzheng Shu 等AAAI 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu 等ACM MM 2023 · 被引用 13 次
- VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video MisinformationYongkang Zhang, Dongyu She, Baiyu Ji, Qichuan Geng 等CVPR 2026
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao 等ACM MM 2021 · 被引用 6 次
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 被引用 28 次
- Joint Geometrical and Statistical Domain Adaptation for Cross-domain Code Vulnerability DetectionQianjin Du, Shiji Zhou, Xiaohui Kuang, Gang Zhao 等EMNLP 2023 · 被引用 4 次
