Coupled Confusion Correction: Learning from Crowds with Sparse Annotations
Hansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan, Shiming Ge
Abstract
As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ab6a7c0-0bdd-4f40-b5fe-6b5e00fbe098Cited by top-tier papers7
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng et al.AAAI 2024 · 63 citations
- Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomTri Nguyen, Shahana Ibrahim, Xiao FuNeurIPS 2024 · 14 citations
- Learning from Noisy Labels via Conditional Distributionally Robust OptimizationHui Guo, Grace Y. Yi, Boyu WangNeurIPS 2024 · 8 citations
- Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex ArgumentsOmar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud PreumEMNLP 2024 · 4 citations
- Collaborative Refining for Learning from Inaccurate LabelsBin Han, Yi-Xuan Sun, Ya-Lin Zhang, Libang Zhang et al.NeurIPS 2024 · 4 citations
Builds on11
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 239 citations
- Do We Need Zero Training Loss After Achieving Zero Training Error?Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu et al.ICML 2020 · 155 citations
- Learning Fast Sample Re-weighting Without Reward DataZizhao Zhang, Tomas PfisterICCV 2021 · 109 citations
Related papers
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
- Improve Learning from Crowds via Generative AugmentationZhendong Chu, Hongning WangKDD 2021 · 7 citations
- Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and RegularizationShahana Ibrahim, Tri Nguyen, Xiao FuICLR 2023 · 3 citations
- On Coresets for End-to-end Learning from CrowdsHang Yang, Zhiwu Li, Witold PedryczAAAI 2026
- SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing LabelsLiyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe et al.AAAI 2026 · 4 citations
