Collaborative Refining for Learning from Inaccurate Labels
Bin Han, Yi-Xuan Sun, Ya-Lin Zhang, Libang Zhang, Haoran Hu, Longfei Li, Jun Zhou, Guo Ye, Huimei He
Abstract
This paper considers the problem of learning from multiple sets of inaccurate labels, which can be easily obtained from low-cost annotators, such as rule-based annotators. Previous works typically concentrate on aggregating information from all the annotators, overlooking the significance of data refinement . This paper presents a collaborative refining approach for learning from inaccurate labels. To refine the data, we introduce the annotator agreement as an instrument, which refers to whether multiple annotators agree or disagree on the labels for a given sample. For samples where some annotators disagree , a comparative strategy is proposed to filter noise. Through theoretical analysis, the correlations among multiple sets of labels, the respective models trained on them, and the true labels are uncovered, so that relatively reliable labels can be identified. For samples where all annotators agree , an aggregating strategy is designed to mitigate potential noise. Guided by theoretical bounds on loss values, a sample selection criterion is introduced and improved to be more robust against potentially problematic values. Through these two modules, all the samples are refined during training, and these refined samples are used to train a lightweight model simultaneously. Extensive experiments are conducted on benchmark and real-world datasets, which demonstrate the superiority of the proposed framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3280a2d-8843-4544-976b-670b2f4359b2Builds on9
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
- Sample Selection with Uncertainty of Losses for Learning with Noisy LabelsXiaobo Xia, Tongliang Liu, Bo Han, Mingming Gong et al.ICLR 2022 · 139 citations
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
- End-to-End Weak SupervisionSalva Rühling Cachay, Benedikt Boecking, Artur DubrawskiNeurIPS 2021 · 48 citations
Related papers
- ULAREF: A Unified Label Refinement Framework for Learning with Inaccurate SupervisionCongyu Qiao, Ning Xu, Yihao Hu, Xin GengICML 2024 · 1 citation
- Coupled-View Deep Classifier Learning from Multiple Noisy AnnotatorsShikun Li, Shiming Ge, Yingying Hua, Chunhui Zhang et al.AAAI 2020 · 30 citations
- Combating Semantic Contamination in Learning with Label NoiseWenxiao Fan, Kan LiAAAI 2025 · 1 citation
- To Aggregate or Not? Learning with Separate Noisy LabelsJiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid et al.KDD 2023 · 21 citations
- Noise Correction on Subjective DatasetsUthman Jinadu, Yi DingACL 2024 · 2 citations
