Candidate Label Set Pruning: A Data-centric Perspective for Deep Partial-label Learning
Shuo He, Chaojie Wang, Guowu Yang, Lei Feng
摘要
Partial-label learning (PLL) allows each training example to be equipped with a set of candidate labels where only one is the true label. Existing deep PLL research focuses on a learning-centric perspective to design various training strategies for label disambiguation i.e., identifying the concealed true label from the candidate label set for model training. However, when the size of the candidate label set becomes excessively large, these learning-centric strategies would be unable to find the true label for model training, thereby causing performance degradation. This motivates us to think from a data-centric perspective and pioneer a new PLLrelated task called candidate label set pruning (CLSP) that aims to filter out certain potential false candidate labels in a training-free manner. To this end, we propose the first CLSP method based on the inconsistency between the representation space and the candidate label space. Specifically, for each candidate label of a training instance, if it is not a candidate label of the instance's nearest neighbors in the representation space, then it has a high probability of being a false label. Based on this intuition, we employ a per-example pruning scheme that filters out a specific proportion of high-probability false candidate labels. Theoretically, we prove an upper bound of the pruning error rate and analyze how the quality of representations affects our proposed method. Empirically, extensive experiments on both benchmark-simulated and real-world PLL datasets validate the great value of CLSP to significantly improve many state-of-the-art deep PLL methods.
Published as a conference paper at ICLR 2024 with various training strategies for label disambiguation in conventional deep PLL research. To this end, we propose the first versatile training-free CLSP method, based on the inconsistency between the representation space and the candidate label space. Specifically, for each candidate label of a training instance, if it is not a candidate label of the instance's nearest neighbors in the representation space, then it has a high probability of being a false label. Based on this intuition, we employ a perexample pruning scheme that filters out a specific proportion of high-probability false candidate labels. Theoretically, we prove an upper bound of the pruning error rate and analyze how the quality of representations affects the proposed algorithm. Empirically, we evaluate the task of CLSP on both benchmark-simulated and real-world datasets across various PLL settings with eleven state-ofthe-art deep PLL methods. Extensive experiments clearly validate the effectiveness of our proposed CLSP method to improve existing PLL methods.
Our main contributions can be summarized as follows:
• A new data-centric perspective for deep PLL. Different from the conventional learning-centric perspective in deep PLL research, we pioneer a new PLL-related task called candidate label set pruning (CLSP) to improve existing deep PLL methods.
• A versatile efficient algorithm. We propose the first versatile training-free CLSP algorithm that prunes a certain proportion of candidates based on the inconsistency between the representation space and candidate label space.
• Theoretical analysis. We theoretically prove an upper bound of the per-example pruning error rate and analyze how the representation quality affects the proposed algorithm.
• Significant experimental improvements. We perform comprehensive experiments on four benchmarks under various PLL settings with eleven state-of-the-art deep PLL methods. Significant improvement validates the superiority of the proposed CLSP method.
Conventional partial-label learning. Early exploration of PLL before the trend of deep learning techniques focused on small-scale datasets with hand-crafted features (Gong et al., 2022). There are mainly two different strategies to handle candidate labels: averaging and identification. The former treats all candidate labels equally (Cour et al., 2011), while the latter aims to identify the concealed true label from candidate labels (Zhang et al., 2016;Xu et al., 2019;Lyu et al., 2019). The drawback of this line of work lies in its limited ability to scale to modern large datasets due to its heavy reliance on hand-crafted features, native linear models, and cost-prohibitive optimization algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsZhaowei Zhu, Jialu Wang, Hao Cheng, Yang LiuICLR 2024 · 被引用 30 次
- Ambiguity-Tolerant Cross-Modal Hashing with Partial LabelsChao Su, Yanan Li, Xu Wang, Yingke Chen 等AAAI 2026 · 被引用 1 次
- Complementary Label Learning with Positive Label Guessing and Negative Label EnhancementYuhang Li, Zhuying Li, Yuheng JiaICLR 2025
- Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label LearningXiaorui Peng, Yuheng Jia, Fuchao Yang, Ran Wang 等ICLR 2025
- Realistic Evaluation of Deep Partial-Label Learning AlgorithmsWei Wang, Dong-Dong Wu, Jindong Wang, Gang Niu 等ICLR 2025
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
相关 Paper
- Revisiting Consistency Regularization for Deep Partial Label LearningDong-Dong Wu, Deng-Bao Wang, Min-Ling ZhangICML 2022 · 被引用 85 次
- Learning with Partial Labels from Semi-supervised PerspectiveXiming Li, Yuanzhi Jiang, Changchun Li, Yiyuan Wang 等AAAI 2023 · 被引用 22 次
- Provably Consistent Partial-Label LearningLei Feng, Jiaqi Lv, Bo Han, Miao Xu 等NeurIPS 2020 · 被引用 188 次
- Partial-label Learning with Mixed Closed-set and Open-set Out-of-candidate ExamplesShuo He, Lei Feng, Guowu YangKDD 2023 · 被引用 1 次
- Semantic Dissimilarity Guided Locality Preserving Projections for Partial Label Dimensionality ReductionYuheng Jia, Jiahao Jiang, Yongheng WangKDD 2023 · 被引用 2 次
