Toward a Perspectivist Turn in Ground Truthing for Predictive Computing
Federico Cabitza, Andrea Campagner, Valerio Basile
Abstract
Most current Artificial Intelligence applications are based on supervised Machine Learning (ML), which ultimately grounds on data annotated by small teams of experts or large ensemble of volunteers. The annotation process is often performed in terms of a majority vote, however this has been proved to be often problematic by recent evaluation studies. In this article, we describe and advocate for a different paradigm, which we call perspectivism: this counters the removal of disagreement and, consequently, the assumption of correctness of traditionally aggregated gold-standard datasets, and proposes the adoption of methods that preserve divergence of opinions and integrate multiple perspectives in the ground truthing process of ML development. Drawing on previous works which inspired it, mainly from the crowdsourcing and multi-rater labeling settings, we survey the state-of-the-art and describe the potential of our proposal for not only the more subjective tasks (e.g. those related to human language) but also those tasks commonly understood as objective (e.g. medical decision making). We present the main benefits of adopting a perspectivist stance in ML, as well as possible disadvantages, and various ways in which such a stance can be implemented in practice. Finally, we share a set of recommendations and outline a research agenda to advance the perspectivist stance in ML.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7833727-5b63-4ac9-9570-55539b98a5c6Cited by top-tier papers28
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 74 citations
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach et al.NeurIPS 2025 · 31 citations
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesBolei Ma, Yuting Li, Wei Zhou, Ziwei Gong et al.ACL 2025 · 28 citations
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke et al.ACL 2023 · 23 citations
- Quantifying the Persona Effect in LLM SimulationsTiancheng Hu, Nigel CollierACL 2024 · 22 citations
Builds on2
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsSangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han et al.CVPR 2021
Related papers
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- PERSEVAL: A Framework for Perspectivist Classification EvaluationSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile et al.EMNLP 2025
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor et al.EMNLP 2025 · 9 citations
- Reliance and Automation for Human-AI Collaborative Data Labeling Conflict ResolutionMichelle Brachman, Zahra Ashktorab, Michael Desmond, Evelyn Duesterwald et al.CSCW 2022 · 16 citations
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
