Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generation
Duc Vu, Anh Nguyen, Chi Tran, Anh Tran
Abstract
Advances in diffusion-based video generation models, while significantly improving human animation, poses threats of misuse through the creation of fake videos from a specific person's photo and text prompts. Recent efforts have focused on adversarial attacks that introduce crafted perturbations to protect images from diffusion-based models. However, most existing approaches target image generation, while relatively few explicitly address image-to-video diffusion models (VDMs), and most primarily focus on UNet-based architectures. Hence, their effectiveness against Diffusion Transformer (DiT) models remains largely under-explored, as these models demonstrate improved feature retention, and stronger temporal consistency due to larger capacity and advanced attention mechanisms. In this work, we introduce Anti-I2V, a novel defense against malicious human image-to-video generation, applicable across diverse diffusion backbones. Instead of restricting noise updates to the RGB space, Anti-I2V operates in both the * and frequency domains, improving robustness and concentrating on salient pixels. We then identify the network layers that capture the most distinct semantic features during the denoising process to design appropriate training objectives that maximize degradation of temporal coherence and generation fidelity. Through extensive validation, Anti-I2V demonstrates state-of-the-art defense performance against diverse video diffusion models, offering an effective solution to the problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 551ea389-ec00-4787-87d7-3b80c915faf3Cited by top-tier papers1
Ask how each one uses itBuilds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- DCT-Shield: A Robust Frequency Domain Defense Against Malicious Image EditingAniruddha Bala, Rohit Chowdhury, Rohan Jaiswal, Siddharth RohedaICCV 2025 · 9 citations
- Anti-DreamBooth: Protecting users from personalized text-to-image synthesisThanh Van Le, Hao Phung, Thuan Hoang Nguyen, Quan Dao et al.ICCV 2023 · 144 citations
- UniDef: Universal Defense Against Unauthorized Image ManipulationMingwen Shao, Lingzhuang Meng, Xiang Lv, Mengyao Wu et al.CVPR 2026
- AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion ModelsYaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang et al.ICML 2025
- PromptFlare: Prompt-Generalized Defense via Cross-Attention Decoy in Diffusion-Based InpaintingHohyun Na, Seunghoo Hong, Simon S. WooACM MM 2025 · 1 citation
