Deep Automodulators
Ari Heljakka, Yuxin Hou, Juho Kannala, Arno Solin
Abstract
We introduce a new category of generative autoencoders called automodulators. These networks can faithfully reproduce individual real-world input images like regular autoencoders, but also generate a fused sample from an arbitrary combination of several such images, allowing instantaneous 'style-mixing' and other new applications. An automodulator decouples the data flow of decoder operations from statistical properties thereof and uses the latent vector to modulate the former by the latter, with a principled approach for mutual disentanglement of decoder layers. Prior work has explored similar decoder architecture with GANs, but their focus has been on random sampling. A corresponding autoencoder could operate on real input images. For the first time, we show how to train such a general-purpose model with sharp outputs in high resolution, using novel training techniques, demonstrated on four image data sets. Besides style-mixing, we show state-of-the-art results in autoencoder comparison, and visual image quality nearly indistinguishable from state-of-the-art GANs. We expect the automodulator variants to become a useful building block for image applications and other data domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46050a23-d824-4598-8604-bb7df2eda0fcCited by top-tier papers2
- Learning Attribute-driven Disentangled Representations for Interactive Fashion RetrievalYuxin Hou, Eleonora Vig, Michael Donoser, Loris BazzaniICCV 2021 · 58 citations
- Diverse Shape Completion via Style Modulated Generative Adversarial NetworksWesley Khademi, Fuxin LiNeurIPS 2023 · 2 citations
Builds on3
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Gaussian Process Priors for View-Aware InferenceYuxin Hou, Ari Heljakka, Arno SolinAAAI 2021 · 1 citation
- Analyzing and Improving the Image Quality of StyleGANTero Karras, Samuli Laine, Miika Aittala, Janne Hellsten et al.CVPR 2020
Related papers
- Adversarial Latent AutoencodersStanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco DorettoCVPR 2020
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
- Towards Controllable and Photorealistic Region-wise Image ManipulationAnsheng You, Chenglin Zhou, Qixuan Zhang, Lan XuACM MM 2021 · 2 citations
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image EditingHyunsu Kim, Yunjey Choi, Junho Kim, Sungjoo Yoo et al.CVPR 2021
