Enhancing Tail Performance in Extreme Classifiers by Label Variance Reduction
Anirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal, Mudit Dhawan, Neelabh Madan, Deepesh Hada, Vidit Jain, Sonu Mehta, Yashoteja Prabhu, Manish Gupta, Ramachandran Ramjee, Manik Varma
Abstract
Extreme Classification (XC) architectures, which utilize a massive One-vs-All (OvA) classifier layer at the output, have demonstrated remarkable performance on problems with large label sets. Nonetheless, these architectures falter on tail labels with few representative samples. This phenomenon has been attributed to factors such as classifier over-fitting and missing label bias, and solutions involving regularization and loss re-calibration have been developed. This paper explores the impact of label variance - a previously unexamined factor - on the tail performance in extreme classifiers. It also develops a method to systematically reduce label variance in XC by transferring the knowledge from a specialized tail-robust teacher model to the OvA classifiers. For this purpose, it proposes a principled knowledge distillation framework, LEVER, which enhances the tail performance in extreme classifiers with formal guarantees on generalization. Comprehensive experiments are conducted on a diverse set of XC datasets, demonstrating that LEVER can enhance tail performance by around 5% and 6% points in PSP and coverage metrics, respectively, when integrated with leading extreme classifiers. Moreover, it establishes a new state-of-the-art when added to the top-performing Renee classifier. Extensive ablations and analyses substantiate the efficacy of our design choices. Another significant contribution is the release of two new XC datasets that are different from and more challenging than the available benchmark datasets, thereby encouraging more rigorous algorithmic evaluation in the future. Code for LEVER is available at: aka.ms/lever.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dce4ecc-f176-4b0b-8f3c-5251b46b6212Cited by top-tier papers4
- Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label FeaturesSiddhant Kharbanda, Devaansh Gupta, Erik Schultheis, Atmadeep Banerjee et al.KDD 2024 · 6 citations
- On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme ClassificationJatin Prakash, Anirudh Buvanesh, Bishal Santra, Deepak Saini et al.KDD 2025 · 1 citation
- Extreme Meta-Classification for Large-Scale Zero-Shot RetrievalSachin Yadav, Deepak Saini, Anirudh Buvanesh, Bhawna Paliwal et al.KDD 2024 · 1 citation
- Navigating Extremes: Dynamic Sparsity in Large Output SpacesNasibullah Nasibullah, Erik Schultheis, Mike Lasby, Yani Ioannou et al.NeurIPS 2024
Builds on12
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang et al.AAAI 2021 · 170 citations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 147 citations
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim et al.ICML 2021 · 97 citations
Related papers
- Convex Surrogates for Unbiased Loss Functions in Extreme Classification With Missing LabelsMohammadreza Qaraei, Erik Schultheis, Priyanshu Gupta, Rohit BabbarWWW 2021 · 29 citations
- Towards Robust Prediction on Tail LabelsTong Wei, Wei-Wei Tu, Yufeng Li, Guo-Ping YangKDD 2021 · 12 citations
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini et al.KDD 2023 · 6 citations
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal et al.WWW 2021 · 71 citations
- Optimizing Tail-Head Trade-off for Extreme Multi-Label Text Classification (XMTC) with RAG-Labels and a Dynamic Two-Stage Retrieval and Fusion PipelineCelso França, Gestefane Rabbi, Thiago Salles, Washington Cunha et al.SIGIR 2025 · 2 citations
