Not All Semantics are Created Equal: Contrastive Self-supervised Learning with Automatic Temperature Individualization
Zi-Hao Qiu, Quanqi Hu, Zhuoning Yuan, Denny Zhou, Lijun Zhang, Tianbao Yang
摘要
In this paper, we aim to optimize a contrastive loss with individualized temperatures in a principled and systematic manner for self-supervised learning. The common practice of using a global temperature parameter τ ignores the fact that "not all semantics are created equal", meaning that different anchor data may have different numbers of samples with similar semantics, especially when data exhibits long-tails. First, we propose a new robust contrastive loss inspired by distributionally robust optimization (DRO), providing us an intuition about the effect of τ and a mechanism for automatic temperature individualization. Then, we propose an efficient stochastic algorithm for optimizing the robust contrastive loss with a provable convergence guarantee without using large minibatch sizes. Theoretical and experimental results show that our algorithm automatically learns a suitable τ for each sample. Specifically, samples with frequent semantics use large temperatures to keep local semantic structures, while samples with rare semantics use small temperatures to induce more separable features. Our method not only outperforms prior strong baselines (e.g., Sim-CLR, CLIP) on unimodal and bimodal datasets with larger improvements on imbalanced data but also is less sensitive to hyper-parameters. To our best knowledge, this is the first methodical approach to optimizing a contrastive loss with individualized temperatures. Most work of Z.H. Qiu was done when visiting the OptMAI lab at TAMU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Learning to Scale Logits for Temperature-Conditional GFlowNetsMinsu Kim, Joohwan Ko, Taeyoung Yun, Dinghuai Zhang 等ICML 2024 · 被引用 31 次
- Contextual Stochastic Bilevel OptimizationYifan Hu, Jie Wang, Yao Xie, Andreas Krause 等NeurIPS 2023 · 被引用 22 次
- To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DROZi-Hao Qiu, Siqi Guo, Mao Xu, Tuo Zhao 等ICML 2024 · 被引用 11 次
- Not All Embeddings are Created Equal: Towards Robust Cross-domain Recommendation via Contrastive LearningWenhao Yang, Yingchun Jian, Yibo Wang, Shiyin Lu 等WWW 2024 · 被引用 10 次
- NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer OptimizationXiyuan Wei, Chih-Jen Lin, Tianbao YangICLR 2026 · 被引用 3 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Temperature Schedules for self-supervised contrastive methods on long-tail dataAnna Kukleva, Moritz Böhle, Bernt Schiele, Hilde Kuehne 等ICLR 2023 · 被引用 7 次
- Understanding Contrastive Learning via Distributionally Robust OptimizationJunkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi 等NeurIPS 2023 · 被引用 55 次
- Provable Stochastic Optimization for Global Contrastive Learning: Small Batch Does Not Harm PerformanceZhuoning Yuan, Yuexin Wu, Zi-Hao Qiu, Xianzhi Du 等ICML 2022 · 被引用 43 次
- On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail LearningJianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao 等ICLR 2023 · 被引用 6 次
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan 等ICLR 2021 · 被引用 296 次
