ReStyle: A Residual-Based StyleGAN Encoder via Iterative Refinement
Yuval Alaluf, Or Patashnik, Daniel Cohen-Or
Abstract
Recently, the power of unconditional image synthesis has significantly advanced through the use of Generative Adversarial Networks (GANs). The task of inverting an image into its corresponding latent code of the trained GAN is of utmost importance as it allows for the manipulation of real images, leveraging the rich semantics learned by the network. Recognizing the limitations of current inversion approaches, in this work we present a novel inversion scheme that extends current encoder-based inversion methods by introducing an iterative refinement mechanism. Instead of directly predicting the latent code of a given real image using a single pass, the encoder is tasked with predicting a residual with respect to the current estimate of the inverted latent code in a self-correcting manner. Our residual-based encoder, named ReStyle, attains improved accuracy compared to current state-of-the-art encoder-based methods with a negligible increase in inference time. We analyze the behavior of ReStyle to gain valuable insights into its iterative nature. We then evaluate the performance of our residual encoder and analyze its robustness compared to optimization-based inversion and state-of-the-art encoders. Code is available via our project page: https: //yuval-alaluf.github.io/restyle-encoder/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f997b200-27f3-40ea-94b8-96dc75c2ac92Cited by top-tier papers96
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano et al.SIGGRAPH 2022 · 501 citations
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
- StyleGAN-XL: Scaling StyleGAN to Large Diverse DatasetsAxel Sauer, Katja Schwarz, Andreas GeigerSIGGRAPH 2022 · 326 citations
- HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image EditingYuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal et al.CVPR 2022 · 250 citations
Builds on20
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
Related papers
- StyleRes: Transforming the Residuals for Real Image Editing with StyleGANHamza Pehlivan, Yusuf Dalva, Aysegul DundarCVPR 2023
- Designing an encoder for StyleGAN image manipulationOmer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik et al.SIGGRAPH 2021 · 692 citations
- HyperInverter: Improving StyleGAN Inversion via HypernetworkTan M. Dinh, Anh Tuan Tran, Rang Nguyen, Binh-Son HuaCVPR 2022 · 111 citations
- RIGID: Recurrent GAN Inversion and Editing of Real Face VideosYangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Ping LuoICCV 2023 · 14 citations
- Dual-path Image Inpainting with Auxiliary GAN InversionWentao Wang, Li Niu, Jianfu Zhang, Xue Yang et al.CVPR 2022 · 46 citations
