Candidate Label Set Pruning: A Data-centric Perspective for Deep Partial-label Learning
Shuo He, Chaojie Wang, Guowu Yang, Lei Feng
Abstract
Partial-label learning (PLL) allows each training example to be equipped with a set of candidate labels where only one is the true label. Existing deep PLL research focuses on a learning-centric perspective to design various training strategies for label disambiguation i.e., identifying the concealed true label from the candidate label set for model training. However, when the size of the candidate label set becomes excessively large, these learning-centric strategies would be unable to find the true label for model training, thereby causing performance degradation. This motivates us to think from a data-centric perspective and pioneer a new PLLrelated task called candidate label set pruning (CLSP) that aims to filter out certain potential false candidate labels in a training-free manner. To this end, we propose the first CLSP method based on the inconsistency between the representation space and the candidate label space. Specifically, for each candidate label of a training instance, if it is not a candidate label of the instance's nearest neighbors in the representation space, then it has a high probability of being a false label. Based on this intuition, we employ a per-example pruning scheme that filters out a specific proportion of high-probability false candidate labels. Theoretically, we prove an upper bound of the pruning error rate and analyze how the quality of representations affects our proposed method. Empirically, extensive experiments on both benchmark-simulated and real-world PLL datasets validate the great value of CLSP to significantly improve many state-of-the-art deep PLL methods.
Published as a conference paper at ICLR 2024 with various training strategies for label disambiguation in conventional deep PLL research. To this end, we propose the first versatile training-free CLSP method, based on the inconsistency between the representation space and the candidate label space. Specifically, for each candidate label of a training instance, if it is not a candidate label of the instance's nearest neighbors in the representation space, then it has a high probability of being a false label. Based on this intuition, we employ a perexample pruning scheme that filters out a specific proportion of high-probability false candidate labels. Theoretically, we prove an upper bound of the pruning error rate and analyze how the quality of representations affects the proposed algorithm. Empirically, we evaluate the task of CLSP on both benchmark-simulated and real-world datasets across various PLL settings with eleven state-ofthe-art deep PLL methods. Extensive experiments clearly validate the effectiveness of our proposed CLSP method to improve existing PLL methods.
Our main contributions can be summarized as follows:
• A new data-centric perspective for deep PLL. Different from the conventional learning-centric perspective in deep PLL research, we pioneer a new PLL-related task called candidate label set pruning (CLSP) to improve existing deep PLL methods.
• A versatile efficient algorithm. We propose the first versatile training-free CLSP algorithm that prunes a certain proportion of candidates based on the inconsistency between the representation space and candidate label space.
• Theoretical analysis. We theoretically prove an upper bound of the per-example pruning error rate and analyze how the representation quality affects the proposed algorithm.
• Significant experimental improvements. We perform comprehensive experiments on four benchmarks under various PLL settings with eleven state-of-the-art deep PLL methods. Significant improvement validates the superiority of the proposed CLSP method.
Conventional partial-label learning. Early exploration of PLL before the trend of deep learning techniques focused on small-scale datasets with hand-crafted features (Gong et al., 2022). There are mainly two different strategies to handle candidate labels: averaging and identification. The former treats all candidate labels equally (Cour et al., 2011), while the latter aims to identify the concealed true label from candidate labels (Zhang et al., 2016;Xu et al., 2019;Lyu et al., 2019). The drawback of this line of work lies in its limited ability to scale to modern large datasets due to its heavy reliance on hand-crafted features, native linear models, and cost-prohibitive optimization algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsZhaowei Zhu, Jialu Wang, Hao Cheng, Yang LiuICLR 2024 · 30 citations
- Ambiguity-Tolerant Cross-Modal Hashing with Partial LabelsChao Su, Yanan Li, Xu Wang, Yingke Chen et al.AAAI 2026 · 1 citation
- Complementary Label Learning with Positive Label Guessing and Negative Label EnhancementYuhang Li, Zhuying Li, Yuheng JiaICLR 2025
- Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label LearningXiaorui Peng, Yuheng Jia, Fuchao Yang, Ran Wang et al.ICLR 2025
- Realistic Evaluation of Deep Partial-Label Learning AlgorithmsWei Wang, Dong-Dong Wu, Jindong Wang, Gang Niu et al.ICLR 2025
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Revisiting Consistency Regularization for Deep Partial Label LearningDong-Dong Wu, Deng-Bao Wang, Min-Ling ZhangICML 2022 · 85 citations
- Learning with Partial Labels from Semi-supervised PerspectiveXiming Li, Yuanzhi Jiang, Changchun Li, Yiyuan Wang et al.AAAI 2023 · 22 citations
- Provably Consistent Partial-Label LearningLei Feng, Jiaqi Lv, Bo Han, Miao Xu et al.NeurIPS 2020 · 188 citations
- Partial-label Learning with Mixed Closed-set and Open-set Out-of-candidate ExamplesShuo He, Lei Feng, Guowu YangKDD 2023 · 1 citation
- Semantic Dissimilarity Guided Locality Preserving Projections for Partial Label Dimensionality ReductionYuheng Jia, Jiahao Jiang, Yongheng WangKDD 2023 · 2 citations
