KFNN: K-Free Nearest Neighbor For Crowdsourcing
Wenjun Zhang, Liangxiao Jiang, Chaoqun Li
Abstract
To reduce annotation costs, it is common in crowdsourcing to collect only a few noisy labels from different crowd workers for each instance. However, the limited noisy labels restrict the performance of label integration algorithms in inferring the unknown true label for the instance. Recent works have shown that leveraging neighbor instances can help alleviate this problem. Yet, these works all assume that each instance has the same neighborhood size, which defies common sense. To address this gap, we propose a novel label integration algorithm called K-free nearest neighbor (KFNN). In KFNN, the neighborhood size of each instance is automatically determined based on its attributes and noisy labels. Specifically, KFNN initially estimates a Mahalanobis distance distribution from the attribute space to model the relationship between each instance and all classes. This distance distribution is then utilized to enhance the multiple noisy label distribution of each instance. Subsequently, a Kalman filter is designed to mitigate the impact of noise incurred by neighbor instances. Finally, KFNN determines the optimal neighborhood size by the max-margin learning. Extensive experimental results demonstrate that KFNN significantly outperforms all the other state-of-the-art algorithms and exhibits greater robustness in various crowdsourcing scenarios. Our codes and datasets are available at https://github.com/jiangliangxiao/KFNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99ccbce3-e1e5-4f8d-b9e0-124b245a9c22Cited by top-tier papers3
- Label Aggregation for Composite Crowd Tasks by Worker Ability Constraint SatisfactionJiyi LiAAAI 2025 · 1 citation
- FedVeer: Self-Adaptive Skew Estimation for Robust Federated LearningYun Xin, Bangqi Pan, Jianfeng Lu, Shuqin Cao et al.ICML 2026
- Towards a Foundation Model for Crowdsourced Label AggregationHao Liu, Jiacheng Liu, Feilong Tang, Long Chen et al.ICLR 2026
Builds on2
Related papers
- Label Distribution Propagation-based Label Completion for CrowdsourcingTong Wu, Liangxiao Jiang, Wenjun Zhang, Chaoqun LiICML 2025
- MAS: Model-Agnostic Active Annotation Strategy for CrowdsourcingWenjun Zhang, Liangxiao Jiang, Chaoqun Li, Shanshan SiICML 2026
- IWBVT: Instance Weighting-based Bias-Variance Trade-off for CrowdsourcingWenjun Zhang, Liangxiao Jiang, Chaoqun LiNeurIPS 2024 · 2 citations
- Unbiased Multi-Label Learning from Crowdsourced AnnotationsMingxuan Xia, Zenan Huang, Runze Wu, Gengyu Lyu et al.ICML 2024 · 1 citation
- Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomTri Nguyen, Shahana Ibrahim, Xiao FuNeurIPS 2024 · 14 citations
