Massively Scaling Heteroscedastic Classifiers
Mark Collier, Rodolphe Jenatton, Basil Mustafa, Neil Houlsby, Jesse Berent, Effrosyni Kokiopoulou
摘要
Heteroscedastic classifiers, which learn a multivariate Gaussian distribution over prediction logits, have been shown to perform well on image classification problems with hundreds to thousands of classes. However, compared to standard classifiers, they introduce extra parameters that scale linearly with the number of classes. This makes them infeasible to apply to larger-scale problems. In addition heteroscedastic classifiers introduce a critical temperature hyperparameter which must be tuned. We propose HET-XL, a heteroscedastic classifier whose parameter count when compared to a standard classifier scales independently of the number of classes. In our large-scale settings, we show that we can remove the need to tune the temperature hyperparameter, by directly learning it on the training data. On large image classification datasets with up to 4B images and 30k classes our method requires 14× fewer additional parameters, does not require tuning the temperature on a held-out set and performs consistently better than the baseline heteroscedastic classifier. HET-XL improves ImageNet 0-shot classification in a multimodal contrastive learning setup which can be viewed as a 3.5 billion class classification problem. *Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pi-DUAL: Using privileged information to distinguish clean from noisy labelsKe Wang, Guillermo Ortiz-Jiménez, Rodolphe Jenatton, Mark Collier 等ICML 2024 · 被引用 7 次
- On the Generalization of Representation Uncertainty in Earth ObservationSpyros Kondylatos, Nikolaos-Ioannis Bountos, Dimitrios Michail, Xiao Xiang Zhu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton 等NeurIPS 2022 · 被引用 359 次
相关 Paper
- When Noisy Labels Meet Long Tail Dilemmas: A Representation Calibration MethodManyi Zhang, Xuyang Zhao, Jun Yao, Chun Yuan 等ICCV 2023 · 被引用 37 次
- SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme ClassificationPranjal Aggarwal, Ameet Deshpande, Karthik R. NarasimhanICML 2023 · 被引用 8 次
- Free Lunch for Few-shot Learning: Distribution CalibrationShuo Yang, Lu Liu, Min XuICLR 2021 · 被引用 378 次
- Correlated Input-Dependent Label Noise in Large-Scale Image ClassificationMark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton 等CVPR 2021
- Theoretical Insights Into Multiclass Classification: A High-dimensional Asymptotic ViewChristos Thrampoulidis, Samet Oymak, Mahdi SoltanolkotabiNeurIPS 2020 · 被引用 46 次
