SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing Labels
Liyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe, Yuta Nakashima
Abstract
Multi-annotator learning (MAL) aims to model annotator-specific labeling patterns. However, existing methods face a critical challenge: they simply skip updating annotator-specific model parameters when encountering missing labels—a common scenario in real-world crowdsourced datasets where each annotator labels only small subsets of samples. This leads to inefficient data utilization and overfitting risks. To this end, we propose a novel similarity-weighted semi-supervised learning framework (SimLabel) that leverages inter-annotator similarities to generate weighted soft labels for missing annotations, enabling the utilization of unannotated samples rather than skipping them entirely. We further introduce a confidence-based iterative refinement mechanism that combines maximum probability with entropy-based uncertainty to prioritize predicted high-quality pseudo-labels to impute missing labels, jointly enhancing similarity estimation and model performance over time. For evaluation, we contribute a new multimodal multi-annotator dataset, AMER2, with high and more variable missing rates, reflecting real-world annotation sparsity and enabling evaluation across different sparsity levels. Extensive experiments validate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 253c83e1-de72-4075-9b10-e1c94cc10796Cited by top-tier papers1
Ask how each one uses itRelated papers
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan et al.AAAI 2024 · 23 citations
- TLLC: Transfer Learning-based Label Completion for CrowdsourcingWenjun Zhang, Liangxiao Jiang, Chaoqun LiICML 2025
- Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label DisambiguationTobias Fuchs, Nadja KleinICML 2026
- All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingIslam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine et al.CVPR 2021
- Beyond Static Alignment: Adaptive Arbitration for Semantic Incongruence in Semi-Supervised Multimodal Sentiment AnalysisHuicong Li, Xiangbo Ji, Wei WuACL 2026
