HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing
Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, Amit Bermano
摘要
The inversion of real images into StyleGAN's latent space is a well-studied problem. Nevertheless, applying existing approaches to real-world scenarios remains an open challenge, due to an inherent trade-off between reconstruction and editability: latent space regions which can accurately represent real images typically suffer from degraded semantic control. Recent work proposes to mitigate this trade-off by fine-tuning the generator to add the target image to well-behaved, editable regions of the latent space. While promising, this fine-tuning scheme is impractical for prevalent use as it requires a lengthy training phase for each new image. In this work, we introduce this approach into the realm of encoder-based inversion. We propose HyperStyle, a hypernetwork that learns to modulate StyleGAN's weights to faithfully express a given image in editable regions of the latent space. A naive modulation approach would require training a hypernetwork with over three billion parameters. Through careful network design, we reduce this to be in line with existing encoders. HyperStyle yields reconstructions comparable to those of optimization techniques with the near real-time inference capabilities of encoders. Lastly, we demonstrate HyperStyle's effectiveness on several applications beyond the inversion task, including the editing of out-of-domain images which were never seen during training. Code is available on our project page: https://yuval-alaluf.github.io/hyperstyle/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper89
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan 等ICLR 2024 · 被引用 223 次
- TF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionShilin Lu, Yanzhu Liu, Adams Wai-Kin KongICCV 2023 · 被引用 214 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 被引用 1,195 次
相关 Paper
- Designing an encoder for StyleGAN image manipulationOmer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik 等SIGGRAPH 2021 · 被引用 692 次
- ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image EditingBingchuan Li, Tianxiang Ma, Peng Zhang, Miao Hua 等AAAI 2023 · 被引用 11 次
- HyperInverter: Improving StyleGAN Inversion via HypernetworkTan M. Dinh, Anh Tuan Tran, Rang Nguyen, Binh-Son HuaCVPR 2022 · 被引用 111 次
- HyperEditor: Achieving Both Authenticity and Cross-Domain Capability in Image Editing via HypernetworksHai Zhang, Chunwei Wu, Guitao Cao, Hailing Wang 等AAAI 2024 · 被引用 6 次
- Diverse Inpainting and Editing with GAN InversionAhmet Burak Yildirim, Hamza Pehlivan, Bahri Batuhan Bilecen, Aysegul DundarICCV 2023 · 被引用 35 次
