SimDA: Simple Diffusion Adapter for Efficient Video Generation
Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, Yu-Gang Jiang
摘要
The recent wave of AI-generated content has witnessed the great development and success of Text-to-Image (T2I) technologies. By contrast, Text-to- Video (T2V) still falls short of expectations though attracting increasing interest. Existing works either train from scratch or adapt large T2I model to videos, both of which are computation and re-source expensive. In this work, we propose a Simple Dif-fusion Adapter (SimDA) that fine-tunes only 24M out of I.IB parameters of a strong T2I model, adapting it to video generation in a parameter-efficient way. In particular, we turn the T2I model for T2V by designing light-weight spatial and temporal adapters for transfer learning. Besides, we change the original spatial attention to the proposed Latent-Shift Attention (LSA) for temporal consistency. With a similar model architecture, we further train a video super-resolution model to generate high-definition (1024 x 1024) videos. In addition to T2V generation in the wild, SimDA could also be utilized in one-shot video editing with only 2 minutes tuning. Doing so, our method could minimize the training effort with extremely few tunable parameters for model adaptation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual GenerationJunke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 132 次
- Human2Robot: Learning Robot Actions from Paired Human-Robot VideosSicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing 等AAAI 2026 · 被引用 15 次
- TAVGBench: Benchmarking Text to Audible-Video GenerationYuxin Mao, Xuyang Shen, Jing Zhang, Zhen Qin 等ACM MM 2024 · 被引用 12 次
- Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and PracticalityTianle Zhang, Langtian Ma, Yuchen Yan, Yuchen Zhang 等NeurIPS 2024 · 被引用 8 次
- Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial FrameQinglong Cao, Xirui Li, Ding Wang, Chao Ma 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper62
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelHan Lin, Jaemin Cho, Abhay Zala, Mohit BansalICLR 2025
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li 等ACM MM 2025 · 被引用 2 次
- LAMP: Learn A Motion Pattern for Few-Shot Video GenerationRuiqi Wu, Liangyu Chen, Tong Yang, Chunle Guo 等CVPR 2024 · 被引用 21 次
