NoisyTwins: Class-Consistent and Diverse Image Generation Through StyleGANs
Harsh Rangwani, Lavish Bansal, Kartik Sharma, Tejan Karmali, Varun Jampani, R. Venkatesh Babu
Abstract
StyleGANs are at the forefront of controllable image generation as they produce a latent space that is semantically disentangled, making it suitable for image editing and manipulation. However, the performance of StyleGANs severely degrades when trained via class-conditioning on large-scale long-tailed datasets. We find that one reason for degradation is the collapse of latents for each class in the W latent space. With NoisyTwins, we first introduce an effective and inexpensive augmentation strategy for class embeddings, which then decorrelates the latents based on self-supervision in the W space. This decorrelation mitigates collapse, ensuring that our method preserves intraclass diversity with class-consistency in image generation. We show the effectiveness of our approach on large-scale real-world long-tailed datasets of ImageNet-LT and iNaturalist 2019, where our method outperforms other methods by ∼ 19% on FID, establishing a new state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f06f4aed-895e-4b5b-a4ae-7ce62fbae7d7Cited by top-tier papers5
- Long-tailed Diffusion Models with Oriented CalibrationTianjiao Zhang, Huangjie Zheng, Jiangchao Yao, Xiangfeng Wang et al.ICLR 2024 · 22 citations
- Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsZhiyao Ren, Yibing Zhan, Liang Ding, Gaoang Wang et al.AAAI 2024 · 15 citations
- DeiT-LT: Distillation Strikes Back for Vision Transformer Training on Long-Tailed DatasetsHarsh Rangwani, Pradipto Mondal, Mayank Mishra, Ashish Ramayee Asokan et al.CVPR 2024 · 12 citations
- Taming the Tail in Class-Conditional GANs: Knowledge Sharing via Unconditional Training at Lower ResolutionsSaeed Khorram, Mingqi Jiang, Mohamad Shahbazi, Mohamad H. Danesh et al.CVPR 2024
- EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion ModelsJingyuan Yang, Jiawei Feng, Hui HuangCVPR 2024
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Diverse Image Generation via Self-Conditioned GANsSteven Liu, Tongzhou Wang, David Bau, Jun-Yan Zhu et al.CVPR 2020
- StyleGAN-XL: Scaling StyleGAN to Large Diverse DatasetsAxel Sauer, Katja Schwarz, Andreas GeigerSIGGRAPH 2022 · 326 citations
- Class-Balancing Diffusion ModelsYiming Qin, Huangjie Zheng, Jiangchao Yao, Mingyuan Zhou et al.CVPR 2023
- Collapse by Conditioning: Training Class-conditional GANs with Limited DataMohamad Shahbazi, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICLR 2022 · 39 citations
- Augmentation-Aware Self-Supervision for Data-Efficient GAN TrainingLiang Hou, Qi Cao, Yige Yuan, Songtao Zhao et al.NeurIPS 2023 · 15 citations
