Aligning Latent and Image Spaces to Connect the Unconnectable
Ivan Skorokhodov, Grigorii Sotnikov, Mohamed Elhoseiny
Abstract
In this work, we develop a method to generate infinite high-resolution images with diverse and complex content. It is based on a perfectly equivariant patch-wise generator with synchronous interpolations in the image and latent spaces. Latent codes, when sampled, are positioned on the coordinate grid, and each pixel is computed from an interpolation of the neighboring codes. We modify the AdaIN mechanism to work in such a setup and train a GAN model to generate images positioned between any two latent vectors. At test time, this allows for generating infinitely large images of diverse scenes that transition naturally from one into another. Apart from that, we introduce LHQ: a new dataset of 90k high-resolution nature landscapes. We test the approach on LHQ, LSUN Tower and LSUN Bridge and outperform the baselines by at least 4 times in terms of quality and diversity of the produced infinite images. The project website is located at https://universome.github.io/alis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be115120-dd6d-4d88-a9c0-7de5f9f63542Cited by top-tier papers35
- Drag Your GAN: Interactive Point-based Manipulation on the Generative Image ManifoldXingang Pan, Ayush Tewari, Thomas Leimkühler, Lingjie Liu et al.SIGGRAPH 2023 · 206 citations
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
- EpiGRAF: Rethinking training of 3D GANsIvan Skorokhodov, Sergey Tulyakov, Yiqun Wang, Peter WonkaNeurIPS 2022 · 145 citations
- Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisWan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen, Yu Cheng et al.AAAI 2023 · 118 citations
- NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual SynthesisJian Liang, Chenfei Wu, Xiaowei Hu, Zhe Gan et al.NeurIPS 2022 · 105 citations
Builds on27
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 933 citations
Related papers
- InfinityGAN: Towards Infinite-Pixel Image SynthesisChieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov et al.ICLR 2022 · 84 citations
- LT3SD: Latent Trees for 3D Scene DiffusionQuan Meng, Lei Li, Matthias Nießner, Angela DaiCVPR 2025
- Progressive Semantic-Aware Style Transformation for Blind Face RestorationChaofeng Chen, Xiaoming Li, Lingbo Yang, Xianhui Lin et al.CVPR 2021
- High-resolution Face Swapping via Latent Semantics DisentanglementYangyang Xu, Bailin Deng, Junle Wang, Yanqing Jing et al.CVPR 2022 · 93 citations
- Patched Denoising Diffusion Models For High-Resolution Image SynthesisZheng Ding, Mengqi Zhang, Jiajun Wu, Zhuowen TuICLR 2024 · 55 citations
