Positional Encoding As Spatial Inductive Bias in GANs
Rui Xu, Xintao Wang, Kai Chen, Bolei Zhou, Chen Change Loy
摘要
SinGAN shows impressive capability in learning internal patch distribution despite its limited effective receptive field. We are interested in knowing how such a translationinvariant convolutional generator could capture the global structure with just a spatially i.i.d. input. In this work, taking SinGAN and StyleGAN2 as examples, we show that such capability, to a large extent, is brought by the implicit positional encoding when using zero padding in the generators. Such positional encoding is indispensable for generating images with high fidelity. The same phenomenon is observed in other generative architectures such as DCGAN and PGGAN. We further show that zero padding leads to an unbalanced spatial bias with a vague relation between locations. To offer a better spatial inductive bias, we investigate alternative positional encodings and analyze their effects. Based on a more flexible positional encoding explicitly, we propose a new multi-scale training strategy and demonstrate its effectiveness in the state-of-the-art unconditional generator StyleGAN2. Besides, the explicit spatial inductive bias substantially improves SinGAN for more versatile image manipulation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- StyTr2: Image Style Transfer with TransformersYingying Deng, Fan Tang, Weiming Dong, Chongyang Ma 等CVPR 2022 · 被引用 345 次
- StyleGAN-XL: Scaling StyleGAN to Large Diverse DatasetsAxel Sauer, Katja Schwarz, Andreas GeigerSIGGRAPH 2022 · 被引用 326 次
- StyleSwin: Transformer-based GAN for High-resolution Image GenerationBowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao 等CVPR 2022 · 被引用 217 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 被引用 933 次
- CARAFE: Content-Aware ReAssembly of FEaturesJiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu 等ICCV 2019 · 被引用 842 次
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan 等ICLR 2021 · 被引用 32 次
相关 Paper
- Toward Spatially Unbiased Generative ModelsJooyoung Choi, Jungbeom Lee, Yonghyun Jeong, Sungroh YoonICCV 2021 · 被引用 17 次
- Analyzing and Improving the Image Quality of StyleGANTero Karras, Samuli Laine, Miika Aittala, Janne Hellsten 等CVPR 2020
- Arbitrary-Scale Image SynthesisEvangelos Ntavelis, Mohamad Shahbazi, Iason Kastanis, Radu Timofte 等CVPR 2022 · 被引用 17 次
- Unveiling The Mask of Position-Information Pattern Through the Mist of Image FeaturesChieh Hubert Lin, Hung-Yu Tseng, Hsin-Ying Lee, Maneesh Kumar Singh 等ICML 2023 · 被引用 3 次
- PetsGAN: Rethinking Priors for Single Image GenerationZicheng Zhang, Yinglu Liu, Congying Han, Hailin Shi 等AAAI 2022 · 被引用 27 次
