From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
Hui Lu, Yi Yu, Song Xia, Yiming Yang, Deepu Rajan, Boon Poh Ng, Alex C. Kot, Xudong Jiang
Abstract
Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowledge of the VFMs to launch potent attacks. This paper investigates a novel and practical adversarial threat scenario: attacking downstream models or MLLMs fine-tuned from open-source VFMs, without requiring access to the victim task, training data, model query, and architecture. In contrast to conventional transfer-based attacks that rely on task-aligned surrogate models, we demonstrate that adversarial vulnerabilities can be exploited directly from the VFMs. To this end, we propose the Transferable Video Attack (TVA), a temporal-aware adversarial attack method that leverages the temporal representation dynamics of VFMs to craft effective perturbations. TVA integrates a bidirectional contrastive learning mechanism to maximize the discrepancy between the clean and adversarial features, and introduces a temporal consistency loss that exploits motion cues to enhance the sequential impact of perturbations. TVA avoids the need to train expensive surrogate models or access to domain-specific data, thereby offering a more practical and efficient attack strategy. Extensive experiments across 24 video-related tasks demonstrate the efficacy of TVA against downstream models and MLLMs, revealing a previously underexplored security vulnerability in the deployment of video models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e64af954-73e4-4307-8624-b3bd3bcb6bd9Cited by top-tier papers4
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi et al.CVPR 2026 · 12 citations
- ApET: Approximation-Error Guided Token Compression for Efficient VLMsQiankun Ma, Ziyao Zhang, Haofei Wang, Zhen Song et al.CVPR 2026 · 11 citations
- Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural NetworksYi Yu, Qixin Zhang, Shuhan Ye, Xun Lin et al.ICLR 2026 · 8 citations
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion ModelsFengpeng Li, Kemou Li, Qizhou Wang, Bo Han et al.ICLR 2026 · 6 citations
Builds on33
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
Related papers
- Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video ApproachLinhao Huang, Xue Jiang, Zhiqiang Wang, Wentao Mo et al.AAAI 2026 · 6 citations
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 48 citations
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 15 citations
- GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu et al.ACM MM 2023 · 13 citations
- Cross-Modal Transferable Adversarial Attacks from Images to VideosZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangCVPR 2022 · 45 citations
