IWBVT: Instance Weighting-based Bias-Variance Trade-off for Crowdsourcing
Wenjun Zhang, Liangxiao Jiang, Chaoqun Li
Abstract
In recent years, a large number of algorithms for label integration and noise correction have been proposed to infer the unknown true labels of instances in crowdsourcing. They have made great advances in improving the label quality of crowdsourced datasets. However, due to the presence of intractable instances, these algorithms are usually not as significant in improving the model quality as they are in improving the label quality. To improve the model quality, this paper proposes an instance weighting-based bias-variance trade-off (IWBVT) approach. IWBVT at first proposes a novel instance weighting method based on the complementary set and entropy, which mitigates the impact of intractable instances and thus makes the bias and variance of trained models closer to the unknown true results. Then, IWBVT performs probabilistic loss regressions based on the bias-variance decomposition, which achieves the bias-variance trade-off and thus reduces the generalization error of trained models. Experimental results indicate that IWBVT can serve as a universal post-processing approach to significantly improving the model quality of existing state-of-the-art label integration algorithms and noise correction algorithms. Our codes and datasets are available at https://github.com
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Label Aggregation for Composite Crowd Tasks by Worker Ability Constraint SatisfactionJiyi LiAAAI 2025 · 1 citation
- Towards a Foundation Model for Crowdsourced Label AggregationHao Liu, Jiacheng Liu, Feilong Tang, Long Chen et al.ICLR 2026
Builds on6
- Learning Cluster-Wise Anchors for Multi-View ClusteringChao Zhang, Xiuyi Jia, Zechao Li, Chunlin Chen et al.AAAI 2024 · 66 citations
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han et al.NeurIPS 2023 · 50 citations
- InstanT: Semi-supervised Learning with Instance-dependent ThresholdsMuyang Li, Runze Wu, Haoyu Liu, Jun Yu et al.NeurIPS 2023 · 27 citations
- Predicting Label Distribution from Multi-label RankingYunan Lu, Xiuyi JiaNeurIPS 2022 · 11 citations
- Generative Calibration of Inaccurate Annotation for Label Distribution LearningLiang He, Yunan Lu, Weiwei Li, Xiuyi JiaAAAI 2024 · 9 citations
Related papers
- Label Correction of Crowdsourced Noisy Annotations with an Instance-Dependent Noise Transition ModelHui Guo, Boyu Wang, Grace YiNeurIPS 2023 · 21 citations
- Label Distribution Propagation-based Label Completion for CrowdsourcingTong Wu, Liangxiao Jiang, Wenjun Zhang, Chaoqun LiICML 2025
- KFNN: K-Free Nearest Neighbor For CrowdsourcingWenjun Zhang, Liangxiao Jiang, Chaoqun LiNeurIPS 2024 · 4 citations
- Resurfacing the Instance-only Dependent Label Noise Model through Loss CorrectionMustafa Enes Aydın, Maarten De Vos, Alexander BertrandICLR 2026
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen et al.ICDE 2023 · 4 citations
