Coupled Confusion Correction: Learning from Crowds with Sparse Annotations
Hansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan, Shiming Ge
摘要
As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng 等AAAI 2024 · 被引用 63 次
- Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomTri Nguyen, Shahana Ibrahim, Xiao FuNeurIPS 2024 · 被引用 14 次
- Learning from Noisy Labels via Conditional Distributionally Robust OptimizationHui Guo, Grace Y. Yi, Boyu WangNeurIPS 2024 · 被引用 8 次
- Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex ArgumentsOmar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud PreumEMNLP 2024 · 被引用 4 次
- Collaborative Refining for Learning from Inaccurate LabelsBin Han, Yi-Xuan Sun, Ya-Lin Zhang, Libang Zhang 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper11
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 被引用 239 次
- Do We Need Zero Training Loss After Achieving Zero Training Error?Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu 等ICML 2020 · 被引用 155 次
- Learning Fast Sample Re-weighting Without Reward DataZizhao Zhang, Tomas PfisterICCV 2021 · 被引用 109 次
相关 Paper
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 被引用 60 次
- Improve Learning from Crowds via Generative AugmentationZhendong Chu, Hongning WangKDD 2021 · 被引用 7 次
- Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and RegularizationShahana Ibrahim, Tri Nguyen, Xiao FuICLR 2023 · 被引用 3 次
- On Coresets for End-to-end Learning from CrowdsHang Yang, Zhiwu Li, Witold PedryczAAAI 2026
- SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing LabelsLiyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe 等AAAI 2026 · 被引用 4 次
