RIGID: Recurrent GAN Inversion and Editing of Real Face Videos
Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Ping Luo
Abstract
GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversion and eDiting (RIGID), to explicitly and simultaneously enforce temporally coherent GAN inversion and facial editing of real videos. Our approach models the temporal relations between current and previous frames from three aspects. To enable a faithful real video reconstruction, we first maximize the inversion fidelity and consistency by learning a temporal compensated latent code. Second, we observe incoherent noises lie in the high-frequency domain that can be disentangled from the latent space. Third, to remove the inconsistency after attribute manipulation, we propose an in-between frame composition constraint such that the arbitrary frame must be a direct composite of its neighboring frames. Our unified framework learns the inherent coherence between input frames in an end-to-end manner, and therefore it is agnostic to a specific attribute and can be applied to arbitrary editing of the same video without re-training. Extensive experiments demonstrate that RIGID outperforms state-of-the-art methods qualitatively and quantitatively in both inversion and editing tasks. The deliverables can be found in https://cnnlstm.github.io/RIGID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dae0d3a-7c3e-4ee6-861e-1597008e8889Cited by top-tier papers3
- Drag Your Noise: Interactive Point-based Editing via Diffusion Semantic PropagationHaofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng et al.CVPR 2024 · 12 citations
- SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingLingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu et al.ACM MM 2024 · 10 citations
- GoodDrag: Towards Good Practices for Drag Editing with Diffusion ModelsZewei Zhang, Huan Liu, Jun Chen, Xiangyu XuICLR 2025
Builds on22
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
Related papers
- From Continuity to Editability: Inverting GANs with Consecutive ImagesYangyang Xu, Yong Du, Wenpeng Xiao, Xuemiao Xu et al.ICCV 2021 · 41 citations
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- InvertAvatar: Incremental GAN Inversion for Generalized Head AvatarsXiaochen Zhao, Jingxiang Sun, Lizhen Wang, Jinli Suo et al.SIGGRAPH 2024 · 10 citations
- StyleRes: Transforming the Residuals for Real Image Editing with StyleGANHamza Pehlivan, Yusuf Dalva, Aysegul DundarCVPR 2023
- Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video EncodingGyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi et al.CVPR 2023
