StyleInV: A Temporal Style Modulated Inversion Network for Unconditional Video Generation
Yuhan Wang, Liming Jiang, Chen Change Loy
摘要
Unconditional video generation is a challenging task that involves synthesizing high-quality videos that are both coherent and of extended duration. To address this challenge, researchers have used pretrained StyleGAN image generators for high-quality frame synthesis and focused on motion generator design. The motion generator is trained in an autoregressive manner using heavy 3D convolutional discriminators to ensure motion coherence during video generation. In this paper, we introduce a novel motion generator design that uses a learning-based inversion network for GAN. The encoder in our method captures rich and smooth priors from encoding images to latents, and given the latent of an initially generated frame as guidance, our method can generate smooth future latent by modulating the inversion encoder temporally. Our method enjoys the advantage of sparse training and naturally constrains the generation space of our motion generator with the inversion network guided by the initial frame, eliminating the need for heavy discriminators. Moreover, our method supports style transfer with simple fine-tuning when the encoder is paired with a pretrained StyleGAN generator. Extensive experiments conducted on various benchmarks demonstrate the superiority of our method in generating long and highresolution videos with decent single-frame quality and temporal consistency. Project website: https://www.mmlabntu.com/project/styleinv/index.html .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Recipe for Scaling up Text-to-Video Generation with Text-free VideosXiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing 等CVPR 2024 · 被引用 19 次
- Using Left and Right Brains Together: Towards Vision and Language PlanningJun Cen, Chenfei Wu, Xiao Liu, Shengming Yin 等ICML 2024 · 被引用 12 次
- CHAIN: Enhancing Generalization in Data-Efficient GANs via LipsCHitz Continuity ConstrAIned NormalizationYao Ni, Piotr KoniuszCVPR 2024 · 被引用 10 次
- Taming Rectified Flow for Inversion and EditingJiangshan Wang, Junfu Pu, Zhongang Qi, Jiayi Guo 等ICML 2025
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
相关 Paper
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 被引用 167 次
- Towards Smooth Video CompositionQihang Zhang, Ceyuan Yang, Yujun Shen, Yinghao Xu 等ICLR 2023 · 被引用 2 次
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 被引用 25 次
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal 等ICCV 2023 · 被引用 3 次
- MoStGAN-V: Video Generation with Temporal Motion StylesXiaoqian Shen, Xiang Li, Mohamed ElhoseinyCVPR 2023
