Convex Surrogates for Unbiased Loss Functions in Extreme Classification With Missing Labels
Mohammadreza Qaraei, Erik Schultheis, Priyanshu Gupta, Rohit Babbar
摘要
Extreme Classification (XC) refers to supervised learning where each training/test instance is labeled with small subset of relevant labels that are chosen from a large set of possible target labels. The framework of XC has been widely employed in web applications such as automatic labeling of web-encyclopedia, prediction of related searches, and recommendation systems. While most state-of-the-art models in XC achieve high overall accuracy by performing well on the frequently occurring labels, they perform poorly on a large number of infrequent (tail) labels. This arises from two statistical challenges, (i) missing labels, as it is virtually impossible to manually assign every relevant label to an instance, and (ii) highly imbalanced data distribution where a large fraction of labels are tail labels. In this work, we consider common loss functions that decompose over labels, and calculate unbiased estimates that compensate missing labels according to Natarajan et al. [26]. This turns out to be disadvantageous from an optimization perspective, as important properties such as convexity and lower-boundedness are lost. To circumvent this problem, we use the fact that typical loss functions in XC are convex surrogates of the 0-1 loss, and thus propose to switch to convex surrogates of its unbiased version. These surrogates are further adapted to the label imbalance by combining with label-frequency-based rebalancing. We show that the proposed loss functions can be easily incorporated into various different frameworks for extreme classification. This includes (i) linear classifiers, such as DiSMEC, on sparse input data representation, (ii) attention-based deep architecture, AttentionXML, learnt on dense Glove embeddings, and (iii) XLNet-based transformer model for extreme classification, APLC-XLNet. Our results demonstrate consistent improvements over the respective vanilla baseline models, on the propensity-scored metrics for precision and nDCG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- CascadeXML: Rethinking Transformers for End-to-end Multi-resolution Training in Extreme Multi-label ClassificationSiddhant Kharbanda, Atmadeep Banerjee, Erik Schultheis, Rohit BabbarNeurIPS 2022 · 被引用 26 次
- Generalized test utilities for long-tail performance in extreme multi-label classificationErik Schultheis, Marek Wydmuch, Wojciech Kotlowski, Rohit Babbar 等NeurIPS 2023 · 被引用 7 次
- Enhancing Tail Performance in Extreme Classifiers by Label Variance ReductionAnirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal 等ICLR 2024 · 被引用 6 次
- InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme ClassificationSiddhant Kharbanda, Atmadeep Banerjee, Devaansh Gupta, Akash Palrecha 等SIGIR 2023 · 被引用 6 次
- Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label FeaturesSiddhant Kharbanda, Devaansh Gupta, Erik Schultheis, Atmadeep Banerjee 等KDD 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 被引用 57 次
- How Well Calibrated are Extreme Multi-label Classifiers? An Empirical AnalysisNasib Ullah, Erik Schultheis, Jinbin Zhang, Rohit BabbarKDD 2025 · 被引用 1 次
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini 等KDD 2023 · 被引用 6 次
- On Missing Labels, Long-tails and Propensities in Extreme Multi-label ClassificationErik Schultheis, Marek Wydmuch, Rohit Babbar, Krzysztof DembczynskiKDD 2022 · 被引用 20 次
- Towards Robust Prediction on Tail LabelsTong Wei, Wei-Wei Tu, Yufeng Li, Guo-Ping YangKDD 2021 · 被引用 12 次
