Learning Semantic-aware Normalization for Generative Adversarial Networks
Heliang Zheng, Jianlong Fu, Yanhong Zeng, Jiebo Luo, Zheng-Jun Zha
Abstract
The recent advances in image generation have been achieved by style-based image generators. Such approaches learn to disentangle latent factors in different image scales and encode latent factors as "style" to control image synthesis. However, existing approaches cannot further disentangle fine-grained semantics from each other, which are often conveyed from feature channels. In this paper, we propose a novel image synthesis approach by learning Semantic-aware relative importance for feature channels in Generative Adversarial Networks (SariGAN). Such a model disentangles latent factors according to the semantic of feature channels by channel-/group-wise fusion of latent codes and feature channels. Particularly, we learn to cluster feature channels by semantics and propose an adaptive group-wise Normalization (AdaGN) to independently control the styles of different channel groups. For example, we can adjust the statistics of channel groups for a human face to control the open and close of the mouth, while keeping other facial features unchanged. We propose to use adversarial training, a channel grouping loss, and a mutual information loss for joint optimization, which not only enables highfidelity image synthesis but leads to superior interpretable properties. Extensive experiments show that our approach outperforms the SOTA style-based approaches in both unconditional image generation and conditional image inpainting tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65fbb3dc-db8e-4bae-9ca9-b9ba77839dbdCited by top-tier papers7
- Diffusion for World Modeling: Visual Details Matter in AtariEloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto et al.NeurIPS 2024 · 359 citations
- Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsHongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun et al.CVPR 2022 · 107 citations
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang et al.NeurIPS 2021 · 31 citations
- PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image ModelsYiming Zhang, Zhening Xing, Yanhong Zeng, Youqing Fang et al.CVPR 2024 · 18 citations
- Contextual Outpainting with Object-Level Contrastive LearningJiacheng Li, Chang Chen, Zhiwei XiongCVPR 2022 · 10 citations
Builds on12
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 933 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
- Controlling generative models with continuous factors of variationsAntoine Plumerault, Hervé Le Borgne, Céline HudelotICLR 2020 · 132 citations
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- Mutual Information Gradient Estimation for Representation LearningLiangjian Wen, Yiji Zhou, Lirong He, Mingyuan Zhou et al.ICLR 2020 · 34 citations
Related papers
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 88 citations
- Diagonal Attention and Style-based GAN for Content-Style Disentanglement in Image Generation and TranslationGihyun Kwon, Jong Chul YeICCV 2021 · 59 citations
- Cluster-guided Image Synthesis with Unconditional ModelsMarkos Georgopoulos, James Oldfield, Grigorios G. Chrysos, Yannis PanagakisCVPR 2022
- Fashion Editing With Adversarial Parsing LearningHaoye Dong, Xiaodan Liang, Yixuan Zhang, Xujie Zhang et al.CVPR 2020
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
