Lune

NeurIPS2022顶会

Understanding Deep Contrastive Learning via Coordinate-wise Optimization

Yuandong Tian

2022年份
51被引次数
21顶会引用

摘要

We show that Contrastive Learning (CL) under a broad family of loss functions (including InfoNCE) has a unified formulation of coordinate-wise optimization on the network parameter θ\boldsymbol{\theta} and pairwise importance α\alpha, where the max player θ\boldsymbol{\theta} learns representation for contrastiveness, and the min player α\alpha puts more weights on pairs of distinct samples that share similar representations. The resulting formulation, called α\alpha-CL, unifies not only various existing contrastive losses, which differ by how sample-pair importance α\alpha is constructed, but also is able to extrapolate to give novel contrastive losses beyond popular ones, opening a new avenue of contrastive loss design. These novel losses yield comparable (or better) performance on CIFAR10, STL-10 and CIFAR-100 than classic InfoNCE. Furthermore, we also analyze the max player in detail: we prove that with fixed α\alpha, max player is equivalent to Principal Component Analysis (PCA) for deep linear network, and almost all local minima are global and rank-1, recovering optimal PCA solutions. Finally, we extend our analysis on max player to 2-layer ReLU networks, showing that its fixed points can have higher ranks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper21

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖