Self-Supervised Video GANs: Learning for Appearance Consistency and Motion Coherency
Sangeek Hyun, Jihwan Kim, Jae-Pil Heo
Abstract
A video can be represented by the composition of appearance and motion. Appearance (or content) expresses the information invariant throughout time, and motion describes the time-variant movement. Here, we propose selfsupervised approaches for video Generative Adversarial Networks (GANs) to achieve the appearance consistency and motion coherency in videos. Specifically, the dual discriminators for image and video individually learn to solve their own pretext tasks; appearance contrastive learning and temporal structure puzzle. The proposed tasks enable the discriminators to learn representations of appearance and temporal context, and force the generator to synthesize videos with consistent appearance and natural flow of motions. Extensive experiments in facial expression and human action public benchmarks show that our method outperforms the state-of-the-art video GANs. Moreover, consistent improvements regardless of the architecture of video GANs confirm that our framework is generic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d87b802-44ee-4c16-851d-a377434c7994Cited by top-tier papers2
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri et al.CVPR 2022 · 36 citations
- RIGID: Recurrent GAN Inversion and Editing of Real Face VideosYangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Ping LuoICCV 2023 · 14 citations
Builds on5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Motion-Based Generator Model: Unsupervised Disentanglement of Appearance, Trackable and Intrackable Motions in Dynamic PatternsJianwen Xie, Ruiqi Gao, Zilong Zheng, Song-Chun Zhu et al.AAAI 2020 · 23 citations
- G3AN: Disentangling Appearance and Motion for Video GenerationYaohui Wang, Piotr Bilinski, François Brémond, Antitza DantchevaCVPR 2020
- Analyzing and Improving the Image Quality of StyleGANTero Karras, Samuli Laine, Miika Aittala, Janne Hellsten et al.CVPR 2020
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj et al.ICCV 2021 · 76 citations
- PV3D: A 3D Generative Model for Portrait Video GenerationEric Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Wenqing Zhang et al.ICLR 2023 · 3 citations
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu et al.ICCV 2021 · 55 citations
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Multiview Pseudo-Labeling for Semi-supervised Learning from VideoBo Xiong, Haoqi Fan, Kristen Grauman, Christoph FeichtenhoferICCV 2021 · 54 citations
