Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
Kento Nozawa, Issei Sato
Abstract
Instance discriminative self-supervised representation learning has been attracted attention thanks to its unsupervised nature and informative feature representation for downstream tasks. In practice, it commonly uses a larger number of negative samples than the number of supervised classes. However, there is an inconsistency in the existing analysis; theoretically, a large number of negative samples degrade classification performance on a downstream supervised task, while empirically, they improve the performance. We provide a novel framework to analyze this empirical result regarding negative samples using the coupon collector's problem. Our bound can implicitly incorporate the supervised loss of the downstream task in the self-supervised loss by increasing the number of negative samples. We confirm that our proposed analysis holds on real-world benchmark datasets. a large number of negative samples are commonly used in self-supervised representation learning algorithms [He et al., 2020 , Chen et al., 2020a]. Contributions. We show difficulty to explain why large negative samples empirically improve supervised accuracy on the downstream task from the CURL framework when we use learned representations as feature vectors for the supervised classification in Section 3. To fill the gap, we propose a novel lower bound to theoretically explain this empirical observation regarding negative samples using the coupon collector's problem in Section 4.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f38a730-2239-42cf-970a-a3d8287849deCited by top-tier papers28
- Understanding Contrastive Learning Requires Incorporating Inductive BiasesNikunj Saunshi, Jordan T. Ash, Surbhi Goel, Dipendra Misra et al.ICML 2022 · 130 citations
- Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation OverlapYifei Wang, Qi Zhang, Yisen Wang, Jiansheng Yang et al.ICLR 2022 · 128 citations
- Augmentations in Graph Contrastive Learning: Current Methodological Flaws & Towards Better PracticesPuja Trivedi, Ekdeep Singh Lubana, Yujun Yan, Yaoqing Yang et al.WWW 2022 · 59 citations
- Do More Negative Samples Necessarily Hurt In Contrastive Learning?Pranjal Awasthi, Nishanth Dikkala, Pritish KamathICML 2022 · 57 citations
- Improving Self-Supervised Learning by Characterizing Idealized RepresentationsYann Dubois, Stefano Ermon, Tatsunori B. Hashimoto, Percy LiangNeurIPS 2022 · 50 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- On the Surrogate Gap between Contrastive and Supervised LossesHan Bao, Yoshihiro Nagano, Kento NozawaICML 2022 · 27 citations
- Incremental False Negative Detection for Contrastive LearningTsai-Shien Chen, Wei-Chih Hung, Hung-Yu Tseng, Shao-Yi Chien et al.ICLR 2022 · 82 citations
- Statistical Consistency and Generalization of Contrastive Representation LearningYuanfan Li, Xiyuan Wei, Tianbao Yang, Yiming YingICML 2026 · 1 citation
- Debiased Contrastive LearningChing-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba et al.NeurIPS 2020 · 761 citations
- Solving Inefficiency of Self-supervised Representation LearningGuangrun Wang, Keze Wang, Guangcong Wang, Philip H. S. Torr et al.ICCV 2021 · 64 citations
