Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks
Alexander Braylan, Omar Alonso, Matthew Lease
摘要
When annotators label data, a key metric for quality assurance is inter-annotator agreement (IAA): the extent to which annotators agree on their labels. Though many IAA measures exist for simple categorical and ordinal labeling tasks, relatively little work has considered more complex labeling tasks, such as structured, multi-object, and free-text annotations. Krippendorff’s α, best known for use with simpler labeling tasks, does have a distance-based formulation with broader applicability, but little work has studied its efficacy and consistency across complex annotation tasks. We investigate the design and evaluation of IAA measures for complex annotation tasks, with evaluation spanning seven diverse tasks: image bounding boxes, image keypoints, text sequence tagging, ranked lists, free text translations, numeric vectors, and syntax trees. We identify the difficulty of interpretability and the complexity of choosing a distance function as key obstacles in applying Krippendorff’s α generally across these tasks. We propose two novel, more interpretable measures, showing they yield more consistent IAA measures across tasks and annotation distance functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Evaluation of African American Language Bias in Natural Language GenerationNicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton 等EMNLP 2023 · 被引用 15 次
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng 等ACL 2025 · 被引用 8 次
- KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision TasksDavid Tschirschwitz, Volker RodehorstCVPR 2026 · 被引用 1 次
它引用的顶会 Paper4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- OpenCrowd: A Human-AI Collaborative Approach for Finding Social Influencers via Open-Ended Answers AggregationInes Arous, Jie Yang, Mourad Khayati, Philippe Cudré-MaurouxWWW 2020 · 被引用 31 次
- Modeling and Aggregation of Complex Annotations via Annotation DistancesAlexander Braylan, Matthew LeaseWWW 2020 · 被引用 15 次
- Aggregating Complex Annotations via Merging and MatchingAlexander Braylan, Matthew LeaseKDD 2021 · 被引用 8 次
相关 Paper
- Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater ReliabilityKa Wong, Praveen K. Paritosh, Lora AroyoACL 2021
- Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation CampaignNicolas Stefanovitch, Jakub PiskorskiEMNLP 2023 · 被引用 1 次
- Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judgeAparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi 等ICLR 2025
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto 等CHI 2021 · 被引用 100 次
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba 等EMNLP 2024 · 被引用 1 次
