Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability
Ka Wong, Praveen K. Paritosh, Lora Aroyo
摘要
When collecting annotations and labeled data from humans, a standard practice is to use inter-rater reliability (IRR) as a measure of data goodness (Hallgren, 2012) . Metrics such as Krippendorff's alpha or Cohen's kappa are typically required to be above a threshold of 0.6 (Landis and Koch, 1977) . These absolute thresholds are unreasonable for crowdsourced data from annotators with high cultural and training variances, especially on subjective topics. We present a new alternative to interpreting IRR that is more empirical and contextualized. It is based upon benchmarking IRR against baseline measures in a replication, one of which is a novel cross-replication reliability (xRR) measure based on Cohen's (1960) kappa. We call this approach the xRR framework. We opensource a replication dataset of 4 million human judgements of facial expressions and analyze it with the proposed framework. We argue this framework can be used to measure the quality of crowdsourced datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- When is Machine Learning Data Good?: Valuing in Public Health DataficationDivy Thakkar, Azra Ismail, Pratyush Kumar, Alex Hanna 等CHI 2022 · 被引用 35 次
- eaSEL: Promoting Social-Emotional Learning and Parent-Child Interaction through AI-Mediated Content ConsumptionJocelyn J. Shen, Jennifer King Chen, Leah Findlater, Griffin Dietz SmithCHI 2025 · 被引用 23 次
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park 等EMNLP 2023 · 被引用 5 次
- Dynamic Multi-Reward Weighting for Multi-Style Controllable GenerationKarin de Langis, Ryan Koo, Dongyeop KangEMNLP 2024 · 被引用 3 次
- D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and EvaluationAida Mostafazadeh Davani, Mark Diaz, Dylan K. Baker, Vinodkumar PrabhakaranEMNLP 2024 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation TasksAlexander Braylan, Omar Alonso, Matthew LeaseWWW 2022 · 被引用 34 次
- FaceExpressions-70k: A Dataset of Perceived Expression DifferencesAvinab Saha, Yu-Chih Chen, Jean-Charles Bazin, Christian Häne 等SIGGRAPH 2025 · 被引用 1 次
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 等ICLR 2020 · 被引用 11 次
- Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial ExpressionsDe'Aira Bryant, Siqi Deng, Nashlie Sephus, Wei Xia 等CVPR 2022 · 被引用 10 次
- MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and BaselinesDávid Javorský, Ondrej Bojar, François YvonACL 2025
