Lune

NeurIPS2022Top-tier venue

Understanding Deep Contrastive Learning via Coordinate-wise Optimization

Yuandong Tian

2022Year
51Citations
21Top-tier citations

Abstract

We show that Contrastive Learning (CL) under a broad family of loss functions (including InfoNCE) has a unified formulation of coordinate-wise optimization on the network parameter θ\boldsymbol{\theta} and pairwise importance α\alpha, where the max player θ\boldsymbol{\theta} learns representation for contrastiveness, and the min player α\alpha puts more weights on pairs of distinct samples that share similar representations. The resulting formulation, called α\alpha-CL, unifies not only various existing contrastive losses, which differ by how sample-pair importance α\alpha is constructed, but also is able to extrapolate to give novel contrastive losses beyond popular ones, opening a new avenue of contrastive loss design. These novel losses yield comparable (or better) performance on CIFAR10, STL-10 and CIFAR-100 than classic InfoNCE. Furthermore, we also analyze the max player in detail: we prove that with fixed α\alpha, max player is equivalent to Principal Component Analysis (PCA) for deep linear network, and almost all local minima are global and rank-1, recovering optimal PCA solutions. Finally, we extend our analysis on max player to 2-layer ReLU networks, showing that its fixed points can have higher ranks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4eadef3c-110a-4e74-9d9e-49a20f5ddcad

Cited by top-tier papers21

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines