Grid Diffusion Models for Text-to-Video Generation
Taegyeong Lee, Soyeong Kwon, Taehwan Kim
摘要
Recent advances in the diffusion models have significantly improved text-to-image generation. However, generating videos from text is a more challenging task than generating images from text, due to the much larger dataset and higher computational cost required. Most existing video generation methods use either a 3D U-Net architecture that considers the temporal dimension or autoregressive generation. These methods require large datasets and are limited in terms of computational costs compared to text-to-image generation. To tackle these challenges, we propose a simple but effective novel grid diffusion for text-to-video generation without temporal dimension in architecture and a large text-video paired dataset. We can generate a high-quality video using a fixed amount of GPU memory regardless of the number of frames by representing the video as a grid image. Additionally, since our method reduces the dimensions of the video to the dimensions of the image, various image-based methods can be applied to videos, such as textguided video manipulation from image manipulation. Our proposed method outperforms the existing methods in both quantitative and qualitative evaluations, demonstrating the suitability of our model for real-world video generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MedSegFactory: Text-Guided Generation of Medical Image-Mask PairsJiawei Mao, Yuhan Wang, Yucheng Tang, Daguang Xu 等ICCV 2025 · 被引用 10 次
- Code2Worlds: Empowering Coding LLMs for 4D World GenerationYi Zhang, Yunshuang Wang, Zeyu Zhang, Hao TangICML 2026 · 被引用 6 次
- SwitchCraft: Training-Free Multi-Event Video Generation with Attention ControlsQianxun Xu, Chenxi Song, Yujun Cai, Chi ZhangCVPR 2026 · 被引用 4 次
- From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute TransitionLing Lo, Kelvin C. K. Chan, Wen-Huang Cheng, Ming-Hsuan YangICCV 2025 · 被引用 3 次
- Identity Preserving 3D Head Stylization with Multiview Score DistillationBahri Batuhan Bilecen, Ahmet Berke Gokmen, Furkan Guzelant, Aysegul DundarICCV 2025 · 被引用 3 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao 等NeurIPS 2024 · 被引用 62 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo 等ICLR 2025
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai 等NeurIPS 2024 · 被引用 48 次
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
