SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution
Chao Wang, Zijin Yang, Yaofei Wang, Yuang Qi, Weiming Zhang, Nenghai Yu, Kejiang Chen
摘要
Recent advancements in video generation technologies have been significant, resulting in their widespread application across multiple domains. However, concerns have been mounting over the potential misuse of generated content. Tracing the origin of generated videos has become crucial to mitigate potential misuse and identify responsible parties. Existing video attribution methods require additional operations or the training of source attribution models, which may degrade video quality or necessitate large amounts of training samples. To address these challenges, we define for the first time the"few-shot training-free generated video attribution"task and propose SWIFT, which is tightly integrated with the temporal characteristics of the video. By leveraging the"Pixel Frames(many) to Latent Frame(one)"temporal mapping within each video chunk, SWIFT applies a fixed-length sliding window to perform two distinct reconstructions: normal and corrupted. The variation in the losses between two reconstructions is then used as an attribution signal. We conducted an extensive evaluation of five state-of-the-art (SOTA) video generation models. Experimental results show that SWIFT achieves over 90% average attribution accuracy with merely 20 video samples across all models and even enables zero-shot attribution for HunyuanVideo, EasyAnimate, and Wan2.2. Our source code is available at https://github.com/wangchao0708/SWIFT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- TALL: Thumbnail Layout for Deepfake Video DetectionYuting Xu, Jian Liang, Gengyun Jia, Ziming Yang 等ICCV 2023 · 被引用 133 次
- Getting ViT in Shape: Scaling Laws for Compute-Optimal Model DesignIbrahim M. Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, Lucas BeyerNeurIPS 2023 · 被引用 122 次
相关 Paper
- Attribution as Retrieval: Model-Agnostic AI-Generated Image AttributionHongsong Wang, Renxi Cheng, Chaolei Han, Jie GuiCVPR 2026 · 被引用 4 次
- Where Did I Come From? Origin Attribution of AI-Generated ImagesZhenting Wang, Chen Chen, Yi Zeng, Lingjuan Lyu 等NeurIPS 2023 · 被引用 44 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
- No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly DetectionZunkai Dai, Ke Li, Jiajia Liu, Jie Yang 等CVPR 2026 · 被引用 6 次
- IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerYuyang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 等CVPR 2025
