G3AN: Disentangling Appearance and Motion for Video Generation
Yaohui Wang, Piotr Bilinski, François Brémond, Antitza Dantcheva
Abstract
Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G 3 AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly outperforms state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion. Source code and pre-trained models are publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ce38317-d37b-46e2-a2ca-7902fe1e5283Cited by top-tier papers25
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen et al.NeurIPS 2023 · 579 citations
- SEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionXinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang et al.ICLR 2024 · 226 citations
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
- CCVS: Context-aware Controllable Video SynthesisGuillaume Le Moing, Jean Ponce, Cordelia SchmidNeurIPS 2021 · 98 citations
Builds on3
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Guided Image-to-Image Translation With Bi-Directional Feature TransformationBadour Albahar, Jia-Bin HuangICCV 2019 · 102 citations
- Learning Fixed Points in Generative Adversarial Networks: From Image-to-Image Translation to Disease Detection and LocalizationMd Mahfuzur Rahman Siddiquee, Zongwei Zhou, Nima Tajbakhsh, Ruibin Feng et al.ICCV 2019 · 97 citations
Related papers
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
- ViSt3D: Video Stylization with 3D CNNAyush Pande, Gaurav SharmaNeurIPS 2023 · 7 citations
- Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetZhimeng Zhang, Lincheng Li, Yu Ding, Changjie FanCVPR 2021
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- PV3D: A 3D Generative Model for Portrait Video GenerationEric Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Wenqing Zhang et al.ICLR 2023 · 3 citations
