Disentangling Sampling and Labeling Bias for Learning in Large-output Spaces
Ankit Singh Rawat, Aditya Krishna Menon, Wittawat Jitkrittum, Sadeep Jayasumana, Felix X. Yu, Sashank J. Reddi, Sanjiv Kumar
Abstract
Negative sampling schemes enable efficient training given a large number of classes, by offering a means to approximate a computationally expensive loss function that takes all labels into account. In this paper, we present a new connection between these schemes and loss modification techniques for countering label imbalance. We show that different negative sampling schemes implicitly trade-off performance on dominant versus rare labels. Further, we provide a unified means to explicitly tackle both sampling bias, arising from working with a subset of all labels, and labeling bias, which is inherent to the data due to label imbalance. We empirically verify our findings on long-tail classification and retrieval benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88ba54fa-d7fa-4735-b70e-7654985bc276Cited by top-tier papers3
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 19 citations
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini et al.KDD 2023 · 6 citations
- On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme ClassificationJatin Prakash, Anirudh Buvanesh, Bishal Santra, Deepak Saini et al.KDD 2025 · 1 citation
Builds on5
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Can We Predict New Facts with Open Knowledge Graph Embeddings? A Benchmark for Open Link PredictionSamuel Broscheit, Kiril Gashteovski, Yanjie Wang, Rainer GemullaACL 2020 · 27 citations
- Extreme Classification via Adversarial Softmax ApproximationRobert Bamler, Stephan MandtICLR 2020 · 25 citations
- Equalization Loss for Long-Tailed Object RecognitionJingru Tan, Changbao Wang, Buyu Li, Quanquan Li et al.CVPR 2020
Related papers
- Supercharging Imbalanced Data Learning With Energy-based Contrastive Representation TransferJunya Chen, Zidi Xiu, Benjamin Goldstein, Ricardo Henao et al.NeurIPS 2021 · 12 citations
- Rethinking the Value of Labels for Improving Class-Imbalanced LearningYuzhe Yang, Zhi XuNeurIPS 2020 · 512 citations
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior PerspectiveZhengzhuo Xu, Zenghao Chai, Chun YuanNeurIPS 2021 · 77 citations
- SSE-SAM: Balancing Head and Tail Classes Gradually Through Stage-Wise SAMXingyu Lyu, Qianqian Xu, Zhiyong Yang, Shaojie Lyu et al.AAAI 2025 · 2 citations
- Long-Tailed Multi-Label Visual Recognition by Collaborative Training on Uniform and Re-Balanced SamplingsHao Guo, Song WangCVPR 2021
