On Missing Labels, Long-tails and Propensities in Extreme Multi-label Classification
Erik Schultheis, Marek Wydmuch, Rohit Babbar, Krzysztof Dembczynski
Abstract
The propensity model introduced by Jain et al. [18] has become a standard approach for dealing with missing and long-tail labels in extreme multi-label classification (XMLC). In this paper, we critically revise this approach showing that despite its theoretical soundness, its application in contemporary XMLC works is debatable. We exhaustively discuss the flaws of the propensity-based approach, and present several recipes, some of them related to solutions used in search engines and recommender systems, that we believe constitute promising alternatives to be followed in XMLC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- CascadeXML: Rethinking Transformers for End-to-end Multi-resolution Training in Extreme Multi-label ClassificationSiddhant Kharbanda, Atmadeep Banerjee, Erik Schultheis, Rohit BabbarNeurIPS 2022 · 26 citations
- Generalized test utilities for long-tail performance in extreme multi-label classificationErik Schultheis, Marek Wydmuch, Wojciech Kotlowski, Rohit Babbar et al.NeurIPS 2023 · 7 citations
- Limited-Supervised Multi-Label Learning with Dependency NoiseYejiang Wang, Yuhai Zhao, Zhengkui Wang, Wen Shan et al.AAAI 2024 · 7 citations
- Enhancing Tail Performance in Extreme Classifiers by Label Variance ReductionAnirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal et al.ICLR 2024 · 6 citations
- InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme ClassificationSiddhant Kharbanda, Atmadeep Banerjee, Devaansh Gupta, Akash Palrecha et al.SIGIR 2023 · 6 citations
Builds on3
- Learning Optimal Tree Models under Beam SearchJingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu et al.ICML 2020 · 72 citations
- SiameseXML: Siamese Networks meet Extreme Classifiers with 100M LabelsKunal Dahiya, Ananye Agarwal, Deepak Saini, Gururaj K et al.ICML 2021 · 61 citations
- Optimal Binary Classification Beyond AccuracyShashank Singh, Justin T. KhimNeurIPS 2022 · 9 citations
Related papers
- How Well Calibrated are Extreme Multi-label Classifiers? An Empirical AnalysisNasib Ullah, Erik Schultheis, Jinbin Zhang, Rohit BabbarKDD 2025 · 1 citation
- Convex Surrogates for Unbiased Loss Functions in Extreme Classification With Missing LabelsMohammadreza Qaraei, Erik Schultheis, Priyanshu Gupta, Rohit BabbarWWW 2021 · 29 citations
- Towards Robust Prediction on Tail LabelsTong Wei, Wei-Wei Tu, Yufeng Li, Guo-Ping YangKDD 2021 · 12 citations
- ELIAS: End-to-End Learning to Index and Search in Large Output SpacesNilesh Gupta, Patrick H. Chen, Hsiang-Fu Yu, Cho-Jui Hsieh et al.NeurIPS 2022 · 19 citations
- Extreme Multi-label Classification from Aggregated LabelsYanyao Shen, Hsiang-Fu Yu, Sujay Sanghavi, Inderjit S. DhillonICML 2020 · 10 citations
