Sampled Estimators For Softmax Must Be Biased
Li-Chung Lin, Yaxu Liu, Chih-Jen Lin
Abstract
Models requiring probabilistic outputs are ubiquitous and used in fields such as natural language processing, contrastive learning, and recommendation systems. The standard method of designing such a model is to output unconstrained logits, which are normalized into probabilities with the softmax function. The normalization involves computing a summation across all classes, which becomes prohibitively expensive for problems with a large number of classes. An important strategy to reduce the cost is to sum over a sampled subset of classes in the softmax function, known as the sampled softmax. It was known that the sampled softmax is biased; the expectation taken over the sampled classes is not equal to the softmax function. Many works focused on reducing the bias by using a better way of sampling the subset. However, while sampled softmax is biased, it is unclear whether an unbiased function different from sampled softmax exists. In this paper, we show that all functions that only access a sampled subset of classes must be biased. With this result, we prevent efforts in finding unbiased loss functions and validate that past efforts devoted to reducing bias are the best we can do.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee683c77-3a27-4fa0-a8ca-494467d06d32Cited by top-tier papers2
- Improved Stochastic Optimization of LogSumExpEgor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel DvurechenskiiICML 2026 · 3 citations
- A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk MinimizationXiyuan Wei, Linli Zhou, Bokun Wang, Chih-Jen Lin et al.ICML 2026
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Debiased Contrastive LearningChing-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba et al.NeurIPS 2020 · 761 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Finite-Sum Coupled Compositional Stochastic Optimization: Theory and ApplicationsBokun Wang, Tianbao YangICML 2022 · 38 citations
Related papers
- A Tale of Two Efficient and Informative Negative Sampling DistributionsShabnam Daghaghi, Tharun Medini, Nicholas Meisburger, Beidi Chen et al.ICML 2021 · 11 citations
- Extreme Classification via Adversarial Softmax ApproximationRobert Bamler, Stephan MandtICLR 2020 · 25 citations
- Rethinking Approximate Gaussian Inference in ClassificationBálint Mucsányi, Nathaël Da Costa, Philipp HennigNeurIPS 2025 · 2 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- On the Independence Assumption in Neurosymbolic LearningEmile van Krieken, Pasquale Minervini, Edoardo M. Ponti, Antonio VergariICML 2024 · 18 citations
