Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability
Ka Wong, Praveen K. Paritosh, Lora Aroyo
Abstract
When collecting annotations and labeled data from humans, a standard practice is to use inter-rater reliability (IRR) as a measure of data goodness (Hallgren, 2012) . Metrics such as Krippendorff's alpha or Cohen's kappa are typically required to be above a threshold of 0.6 (Landis and Koch, 1977) . These absolute thresholds are unreasonable for crowdsourced data from annotators with high cultural and training variances, especially on subjective topics. We present a new alternative to interpreting IRR that is more empirical and contextualized. It is based upon benchmarking IRR against baseline measures in a replication, one of which is a novel cross-replication reliability (xRR) measure based on Cohen's (1960) kappa. We call this approach the xRR framework. We opensource a replication dataset of 4 million human judgements of facial expressions and analyze it with the proposed framework. We argue this framework can be used to measure the quality of crowdsourced datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24be0f12-fd23-4ff3-b59d-96e93aa57e1bCited by top-tier papers7
- When is Machine Learning Data Good?: Valuing in Public Health DataficationDivy Thakkar, Azra Ismail, Pratyush Kumar, Alex Hanna et al.CHI 2022 · 35 citations
- eaSEL: Promoting Social-Emotional Learning and Parent-Child Interaction through AI-Mediated Content ConsumptionJocelyn J. Shen, Jennifer King Chen, Leah Findlater, Griffin Dietz SmithCHI 2025 · 23 citations
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park et al.EMNLP 2023 · 5 citations
- Dynamic Multi-Reward Weighting for Multi-Style Controllable GenerationKarin de Langis, Ryan Koo, Dongyeop KangEMNLP 2024 · 3 citations
- D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and EvaluationAida Mostafazadeh Davani, Mark Diaz, Dylan K. Baker, Vinodkumar PrabhakaranEMNLP 2024 · 3 citations
Builds on1
Related papers
- Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation TasksAlexander Braylan, Omar Alonso, Matthew LeaseWWW 2022 · 34 citations
- FaceExpressions-70k: A Dataset of Perceived Expression DifferencesAvinab Saha, Yu-Chih Chen, Jean-Charles Bazin, Christian Häne et al.SIGGRAPH 2025 · 1 citation
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial ExpressionsDe'Aira Bryant, Siqi Deng, Nashlie Sephus, Wei Xia et al.CVPR 2022 · 10 citations
- MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and BaselinesDávid Javorský, Ondrej Bojar, François YvonACL 2025
