Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification
Andreas Grivas, Antonio Vergari, Adam Lopez
摘要
Sigmoid output layers are widely used in multi-label classification (MLC) tasks, in which multiple labels can be assigned to any input. In many practical MLC tasks, the number of possible labels is in the thousands, often exceeding the number of input features and resulting in a low-rank output layer. In multi-class classification, it is known that such a lowrank output layer is a bottleneck that can result in unargmaxable classes: classes which cannot be predicted for any input. In this paper, we show that for MLC tasks, the analogous sigmoid bottleneck results in exponentially many unargmaxable label combinations. We explain how to detect these unargmaxable outputs and demonstrate their presence in three widely used MLC datasets. We then show that they can be prevented in practice by introducing a Discrete Fourier Transform (DFT) output layer, which guarantees that all sparse label combinations with up to k active labels are argmaxable. Our DFT layer trains faster and is more parameter efficient, matching the F1@k score of a sigmoid layer while using up to 50% fewer trainable parameters. Our code is publicly available at https://github.com/andreasgrv/sigmoid-bottleneck . 1 Adding a bias term to the BSL allows the hyperplanes to have offsets and they will not necessarily meet at the origin. However, this cannot solve the problem: such a BSL is more restricted than increasing d by one: we still only get 7 out of 8 label combinations. 2 There are also unargmaxable test examples for d = 100 and d = 200, we chose this example as it had fewer active labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On the Theoretical Limitations of Embedding-Based RetrievalOrion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk LeeICLR 2026 · 被引用 138 次
- On the Independence Assumption in Neurosymbolic LearningEmile van Krieken, Pasquale Minervini, Edoardo M. Ponti, Antonio VergariICML 2024 · 被引用 18 次
- The Theory and Practice of MAP Inference over Non-Convex ConstraintsLeander Kurscheidt, Gabriele Masina, Roberto Sebastiani, Antonio VergariICML 2026 · 被引用 1 次
它引用的顶会 Paper2
- Node Embeddings and Exact Low-Rank Representations of Complex NetworksSudhanshu Chanpuriya, Cameron Musco, Konstantinos Sotiropoulos, Charalampos E. TsourakakisNeurIPS 2020 · 被引用 41 次
- Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in PracticeAndreas Grivas, Nikolay Bogoychev, Adam LopezACL 2022
相关 Paper
- Multilabel Classification by Hierarchical Partitioning and Data-dependent GroupingShashanka Ubaru, Sanjeeb Dash, Arya Mazumdar, Oktay GünlükNeurIPS 2020 · 被引用 10 次
- Reducing Information Bottleneck for Weakly Supervised Semantic SegmentationJungbeom Lee, Jooyoung Choi, Jisoo Mok, Sungroh YoonNeurIPS 2021 · 被引用 174 次
- Plastic Learning with Deep Fourier FeaturesAlex Lewandowski, Dale Schuurmans, Marlos C. MachadoICLR 2025
- Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based ApproachZixun WangCVPR 2026 · 被引用 1 次
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 被引用 14 次
