Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks
Alexander Braylan, Omar Alonso, Matthew Lease
Abstract
When annotators label data, a key metric for quality assurance is inter-annotator agreement (IAA): the extent to which annotators agree on their labels. Though many IAA measures exist for simple categorical and ordinal labeling tasks, relatively little work has considered more complex labeling tasks, such as structured, multi-object, and free-text annotations. Krippendorff’s α, best known for use with simpler labeling tasks, does have a distance-based formulation with broader applicability, but little work has studied its efficacy and consistency across complex annotation tasks. We investigate the design and evaluation of IAA measures for complex annotation tasks, with evaluation spanning seven diverse tasks: image bounding boxes, image keypoints, text sequence tagging, ranked lists, free text translations, numeric vectors, and syntax trees. We identify the difficulty of interpretability and the complexity of choosing a distance function as key obstacles in applying Krippendorff’s α generally across these tasks. We propose two novel, more interpretable measures, showing they yield more consistent IAA measures across tasks and annotation distance functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d3e2bca-7ea8-489d-837b-adc2499156deCited by top-tier papers3
- Evaluation of African American Language Bias in Natural Language GenerationNicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton et al.EMNLP 2023 · 15 citations
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng et al.ACL 2025 · 8 citations
- KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision TasksDavid Tschirschwitz, Volker RodehorstCVPR 2026 · 1 citation
Builds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- OpenCrowd: A Human-AI Collaborative Approach for Finding Social Influencers via Open-Ended Answers AggregationInes Arous, Jie Yang, Mourad Khayati, Philippe Cudré-MaurouxWWW 2020 · 31 citations
- Modeling and Aggregation of Complex Annotations via Annotation DistancesAlexander Braylan, Matthew LeaseWWW 2020 · 15 citations
- Aggregating Complex Annotations via Merging and MatchingAlexander Braylan, Matthew LeaseKDD 2021 · 8 citations
Related papers
- Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater ReliabilityKa Wong, Praveen K. Paritosh, Lora AroyoACL 2021
- Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation CampaignNicolas Stefanovitch, Jakub PiskorskiEMNLP 2023 · 1 citation
- Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judgeAparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi et al.ICLR 2025
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto et al.CHI 2021 · 100 citations
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
