Learning to See by Looking at Noise
Manel Baradad Jurjo, Jonas Wulff, Tongzhou Wang, Phillip Isola, Antonio Torralba
Abstract
Current vision systems are trained on huge datasets, and these datasets come with costs: curation is expensive, they inherit human biases, and there are concerns over privacy and usage rights. To counter these costs, interest has surged in learning from cheaper data sources, such as unlabeled images. In this paper we go a step further and ask if we can do away with real image datasets entirely, instead learning from noise processes. We investigate a suite of image generation models that produce images from simple random processes. These are then used as training data for a visual representation learner with a contrastive loss. We study two types of noise processes, statistical image models and deep generative models under different random initializations. Our findings show that it is important for the noise to capture certain structural properties of real data but that good performance can be achieved even with processes that are far from realistic. We also find that diversity is a key property to learn good representations. Datasets, models, and code are available at https://mbaradad.github.io/learning_with_noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d5e28f0-279e-4f4b-b139-5cf453b92a0cCited by top-tier papers36
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang et al.NeurIPS 2023 · 251 citations
- Generative Models as a Data Source for Multiview Representation LearningAli Jahanian, Xavier Puig, Yonglong Tian, Phillip IsolaICLR 2022 · 148 citations
- Virtual Homogeneity Learning: Defending against Data Heterogeneity in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xin He et al.ICML 2022 · 105 citations
- FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsLihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi et al.NeurIPS 2023 · 94 citations
- PixMix: Dreamlike Pictures Comprehensively Improve Safety MeasuresDan Hendrycks, Andy Zou, Mantas Mazeika, Leonard Tang et al.CVPR 2022 · 93 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 933 citations
- A critical analysis of self-supervision, or what we can learn from a single imageYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 152 citations
Related papers
- Which is Better for Learning with Noisy Labels: The Semi-supervised Method or Modeling Label Noise?Yu Yao, Mingming Gong, Yuxuan Du, Jun Yu et al.ICML 2023 · 15 citations
- Self-Calibrated Variance-Stabilizing Transformations for Real-World Image DenoisingSébastien Herbreteau, Michael UnserICCV 2025 · 3 citations
- Noisier2Noise: Learning to Denoise From Unpaired Noisy DataNick Moran, Dan Schmidt, Yu Zhong, Patrick CoadyCVPR 2020
- Procedural Image Programs for Representation LearningManel Baradad, Chun-Fu Richard Chen, Jonas Wulff, Tongzhou Wang et al.NeurIPS 2022 · 40 citations
- C2N: Practical Generative Noise Modeling for Real-World DenoisingGeonwoon Jang, Wooseok Lee, Sanghyun Son, Kyoung Mu LeeICCV 2021 · 109 citations
