Towards Smooth Video Composition
Qihang Zhang, Ceyuan Yang, Yujun Shen, Yinghao Xu, Bolei Zhou
Abstract
Video generation requires synthesizing consistent and persistent frames with dynamic content over time. This work investigates modeling the temporal relations for composing video with arbitrary length, from a few frames to even infinite, using generative adversarial networks (GANs). First, towards composing adjacent frames, we show that the alias-free operation for single image generation, together with adequately pre-learned knowledge, brings a smooth frame transition without compromising the per-frame quality. Second, by incorporating the temporal shift module (TSM), originally designed for video understanding, into the discriminator, we manage to advance the generator in synthesizing more consistent dynamics. Third, we develop a novel B-Spline based motion representation to ensure temporal smoothness to achieve infinite-length video generation. It can go beyond the frame number used in training. A low-rank temporal modulation is also proposed to alleviate repeating contents for long video generation. We evaluate our approach on various datasets and show substantial improvements over video generation baselines. Code and models will be publicly available at https://genforce.github.io/StyleSV .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ae11afd-0aaa-404f-9bcd-1991925ebd73Cited by top-tier papers4
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 370 citations
- StyleInV: A Temporal Style Modulated Inversion Network for Unconditional Video GenerationYuhan Wang, Liming Jiang, Chen Change LoyICCV 2023 · 20 citations
- Learning Modulated Transformation in GANsCeyuan Yang, Qihang Zhang, Yinghao Xu, Jiapeng Zhu et al.NeurIPS 2023 · 1 citation
- VideoGigaGAN: Towards Detail-rich Video Super-ResolutionYiran Xu, Taesung Park, Richard Zhang, Yang Zhou et al.CVPR 2025
Builds on20
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
Related papers
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal et al.ICCV 2023 · 3 citations
- MoStGAN-V: Video Generation with Temporal Motion StylesXiaoqian Shen, Xiang Li, Mohamed ElhoseinyCVPR 2023
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
- Generating Videos with Dynamics-aware Implicit Generative Adversarial NetworksSihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim et al.ICLR 2022 · 227 citations
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
