Towards Robust Prediction on Tail Labels
Tong Wei, Wei-Wei Tu, Yufeng Li, Guo-Ping Yang
Abstract
Extreme multi-label learning (XML) works to annotate objects with relevant labels from an extremely large label set. Many previous methods treat labels uniformly such that the learned model tends to perform better on head labels, while the performance is severely deteriorated for tail labels. However, it is often desirable to predict more tail labels in many real-world applications. To alleviate this problem, in this work, we show theoretical and experimental evidence for the inferior performance of representative XML methods on tail labels. Our finding is that the norm of label classifier weights typically follows a long-tailed distribution similar to the label frequency, which results in the over-suppression of tail labels. Base on this new finding, we present two new modules: (1)ReRank works to re-rank the predicted score, which significantly improves the performance on tail labels by eliminating the effect of label-priors; (2)Taug augments tail labels via a decoupled learning scheme, which can yield more balanced classification boundary. We conduct experiments on commonly used XML benchmarks with hundreds of thousands of labels, showing that the proposed methods improve the performance of many state-of-the-art XML models by a considerable margin (6% performance gain with respect to [email protected] on average). Anonymous source code is available at https://github.com/ReRANK-XML/rerank-XML.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8877be2-3c05-461f-b0dd-24fdedaa6560Cited by top-tier papers6
- Long-Tail Learning with Foundation Model: Heavy Fine-Tuning HurtsJiang-Xin Shi, Tong Wei, Zhi Zhou, Jie-Jing Shao et al.ICML 2024 · 78 citations
- Metadata-Induced Contrastive Learning for Zero-Shot Multi-Label Text ClassificationYu Zhang, Zhihong Shen, Chieh-Han Wu, Boya Xie et al.WWW 2022 · 34 citations
- Weakly Supervised Multi-Label Classification of Full-Text Scientific PapersYu Zhang, Bowen Jin, Xiusi Chen, Yanzhen Shen et al.KDD 2023 · 8 citations
- Generalized test utilities for long-tail performance in extreme multi-label classificationErik Schultheis, Marek Wydmuch, Wojciech Kotlowski, Rohit Babbar et al.NeurIPS 2023 · 7 citations
- Enhancing Tail Performance in Extreme Classifiers by Label Variance ReductionAnirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal et al.ICLR 2024 · 6 citations
Builds on1
Related papers
- Long-tailed Recognition with Model RebalancingJiaan Luo, Feng Hong, Qiang Hu, Xiaofeng Cao et al.NeurIPS 2025 · 12 citations
- How Well Calibrated are Extreme Multi-label Classifiers? An Empirical AnalysisNasib Ullah, Erik Schultheis, Jinbin Zhang, Rohit BabbarKDD 2025 · 1 citation
- SiameseXML: Siamese Networks meet Extreme Classifiers with 100M LabelsKunal Dahiya, Ananye Agarwal, Deepak Saini, Gururaj K et al.ICML 2021 · 61 citations
- Long-Tailed Partial Label Learning by Head Classifier and Tail Classifier CooperationYuheng Jia, Xiaorui Peng, Ran Wang, Min-Ling ZhangAAAI 2024 · 21 citations
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang et al.AAAI 2021 · 170 citations
