Extreme Classification via Adversarial Softmax Approximation
Robert Bamler, Stephan Mandt
摘要
Training a classifier over a large number of classes, known as 'extreme classification', has become a topic of major interest with applications in technology, science, and e-commerce. Traditional softmax regression induces a gradient cost proportional to the number of classes , which often is prohibitively expensive. A popular scalable softmax approximation relies on uniform negative sampling, which suffers from slow convergence due a poor signal-to-noise ratio. In this paper, we propose a simple training method for drastically enhancing the gradient signal by drawing negative samples from an adversarial model that mimics the data distribution. Our contributions are three-fold: (i) an adversarial sampling mechanism that produces negative samples at a cost only logarithmic in , thus still resulting in cheap gradient updates; (ii) a mathematical proof that this adversarial sampling minimizes the gradient variance while any bias due to non-uniform sampling can be removed; (iii) experimental results on large scale data sets that show a reduction of the training time by an order of magnitude relative to several competitive baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Neural Transformation Learning for Deep Anomaly Detection Beyond ImagesChen Qiu, Timo Pfrommer, Marius Kloft, Stephan Mandt 等ICML 2021 · 被引用 171 次
- Convex Surrogates for Unbiased Loss Functions in Extreme Classification With Missing LabelsMohammadreza Qaraei, Erik Schultheis, Priyanshu Gupta, Rohit BabbarWWW 2021 · 被引用 29 次
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 被引用 19 次
- Disentangling Sampling and Labeling Bias for Learning in Large-output SpacesAnkit Singh Rawat, Aditya Krishna Menon, Wittawat Jitkrittum, Sadeep Jayasumana 等ICML 2021 · 被引用 13 次
- A Tale of Two Efficient and Informative Negative Sampling DistributionsShabnam Daghaghi, Tharun Medini, Nicholas Meisburger, Beidi Chen 等ICML 2021 · 被引用 11 次
相关 Paper
- ANN Softmax: Acceleration of Extreme Classification TrainingKang Zhao, Liuyihan Song, Yingya Zhang, Pan Pan 等VLDB 2022 · 被引用 8 次
- SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme ClassificationPranjal Aggarwal, Ameet Deshpande, Karthik R. NarasimhanICML 2023 · 被引用 8 次
- Sampled Estimators For Softmax Must Be BiasedLi-Chung Lin, Yaxu Liu, Chih-Jen LinNeurIPS 2025 · 被引用 3 次
- Multiclass Boosting and the Cost of Weak LearningNataly Brukhim, Elad Hazan, Shay Moran, Indraneel Mukherjee 等NeurIPS 2021 · 被引用 16 次
- Generative Adversarial Minority OversamplingSankha Subhra Mullick, Shounak Datta, Swagatam DasICCV 2019 · 被引用 222 次
