Let the Prototype Guide You: Robust Aggregation of Sparse Multi-Class Annotations via Annotator Prototype Learning
Ju Chen, Jun Feng, Shenyu Zhang
Abstract
Truth inference is a critical technique for aggregating noisy and biased multi-class classification annotations. State-of-the-art approaches model each annotator using an individual confusion matrix. While well-grounded, they suffer from two fundamental bottlenecks: 1) confusion matrices are underfit when annotators label only a small subset of tasks or when classes are imbalanced, and 2) a single confusion matrix per annotator is inadequate for capturing complex annotator behaviors, leading to class-level collapse when tasks are extremely difficult. Simultaneously addressing these challenges is non-trivial, as it demands both robustness to data sparsity and sufficient expressiveness for complex annotator patterns. In this paper, we propose CPBCC (Class-specific Prototype-driven Bayesian Classifier Combination), which creatively models annotators through a dual-pathway architecture: (i) learning class-specific prototype annotation patterns across all annotators, and (ii) learning annotator-specific weights over prototypes. This framework addresses the bottlenecks and achieves a robust yet rich annotator characterization. Experiments across 10 real-world datasets spanning five domains demonstrate that CPBCC yields a 26% accuracy improvement in the best case, and boosts average accuracy from 68.73% to 74.11%. Our source code is available at https://github.com/JuJuCHEN-HHU/CPBCC_PTBCC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d88d202f-9684-4f0f-ad7e-c01e441b5c3cBuilds on11
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
- Open Knowledge Enrichment for Long-tail EntitiesErmei Cao, Difeng Wang, Jiacheng Huang, Wei HuWWW 2020 · 51 citations
- Adversarial Learning from CrowdsPengpeng Chen, Hailong Sun, Yongqiang Yang, Zhijun ChenAAAI 2022 · 16 citations
- Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix FactorizationShahana Ibrahim, Xiao FuICML 2021 · 12 citations
- DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph LearningFan Li, Xiaoyang Wang, Wenjie Zhang, Ying Zhang et al.ICLR 2026 · 9 citations
Related papers
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan et al.AAAI 2024 · 23 citations
- Coupled-View Deep Classifier Learning from Multiple Noisy AnnotatorsShikun Li, Shiming Ge, Yingying Hua, Chunhui Zhang et al.AAAI 2020 · 30 citations
- Aggregating Complex Annotations via Merging and MatchingAlexander Braylan, Matthew LeaseKDD 2021 · 8 citations
- QuMAB: Query-based Multi-annotator Behavior Pattern LearningLiyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe et al.AAAI 2026 · 3 citations
- Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomTri Nguyen, Shahana Ibrahim, Xiao FuNeurIPS 2024 · 14 citations
