Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification
Andreas Grivas, Antonio Vergari, Adam Lopez
Abstract
Sigmoid output layers are widely used in multi-label classification (MLC) tasks, in which multiple labels can be assigned to any input. In many practical MLC tasks, the number of possible labels is in the thousands, often exceeding the number of input features and resulting in a low-rank output layer. In multi-class classification, it is known that such a lowrank output layer is a bottleneck that can result in unargmaxable classes: classes which cannot be predicted for any input. In this paper, we show that for MLC tasks, the analogous sigmoid bottleneck results in exponentially many unargmaxable label combinations. We explain how to detect these unargmaxable outputs and demonstrate their presence in three widely used MLC datasets. We then show that they can be prevented in practice by introducing a Discrete Fourier Transform (DFT) output layer, which guarantees that all sparse label combinations with up to k active labels are argmaxable. Our DFT layer trains faster and is more parameter efficient, matching the F1@k score of a sigmoid layer while using up to 50% fewer trainable parameters. Our code is publicly available at https://github.com/andreasgrv/sigmoid-bottleneck . 1 Adding a bias term to the BSL allows the hyperplanes to have offsets and they will not necessarily meet at the origin. However, this cannot solve the problem: such a BSL is more restricted than increasing d by one: we still only get 7 out of 8 label combinations. 2 There are also unargmaxable test examples for d = 100 and d = 200, we chose this example as it had fewer active labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5d1d72c-a59d-449c-a5fa-d20f9fd614c7Cited by top-tier papers3
- On the Theoretical Limitations of Embedding-Based RetrievalOrion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk LeeICLR 2026 · 138 citations
- On the Independence Assumption in Neurosymbolic LearningEmile van Krieken, Pasquale Minervini, Edoardo M. Ponti, Antonio VergariICML 2024 · 18 citations
- The Theory and Practice of MAP Inference over Non-Convex ConstraintsLeander Kurscheidt, Gabriele Masina, Roberto Sebastiani, Antonio VergariICML 2026 · 1 citation
Builds on2
- Node Embeddings and Exact Low-Rank Representations of Complex NetworksSudhanshu Chanpuriya, Cameron Musco, Konstantinos Sotiropoulos, Charalampos E. TsourakakisNeurIPS 2020 · 41 citations
- Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in PracticeAndreas Grivas, Nikolay Bogoychev, Adam LopezACL 2022
Related papers
- Multilabel Classification by Hierarchical Partitioning and Data-dependent GroupingShashanka Ubaru, Sanjeeb Dash, Arya Mazumdar, Oktay GünlükNeurIPS 2020 · 10 citations
- Reducing Information Bottleneck for Weakly Supervised Semantic SegmentationJungbeom Lee, Jooyoung Choi, Jisoo Mok, Sungroh YoonNeurIPS 2021 · 174 citations
- Plastic Learning with Deep Fourier FeaturesAlex Lewandowski, Dale Schuurmans, Marlos C. MachadoICLR 2025
- Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based ApproachZixun WangCVPR 2026 · 1 citation
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 14 citations
