PERSEVAL: A Framework for Perspectivist Classification Evaluation
Soda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile, Franco Sansonetti, Antonio Uva, Davide Bernardi
Abstract
Data perspectivism goes beyond majority vote label aggregation by recognizing various perspectives as legitimate ground truths. However, current evaluation practices remain fragmented, making it difficult to compare perspectivist approaches and analyze their impact on different users and demographic subgroups. To address this gap, we introduce PERSEVAL, the first unified framework for evaluating perspectivist models in NLP. A key innovation is its evaluation at the individual annotator level and its treatment of annotators and users as distinct entities, consistently with real-world scenarios. We demonstrate PERSEVAL's capabilities through experiments with both Encoderbased and Decoder-based approaches, as well as an analysis of the effect of sociodemographic prompting. By considering global, text-, traitand user-level evaluation metrics, we show that PERSEVAL is a powerful tool for examining how models are influenced by user-specific information and identifying the biases this information may introduce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09dfafd7-019e-4514-b98e-f9ce45a28c0aBuilds on13
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 236 citations
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 74 citations
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano et al.ACL 2025 · 34 citations
- EPIC: Multi-Perspective Annotation of a Corpus of IronySimona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo et al.ACL 2023 · 9 citations
Related papers
- MultiPICo: Multilingual Perspectivist Irony CorpusSilvia Casola, Simona Frenda, Soda Marem Lo, Erhan Sezerer et al.ACL 2024 · 2 citations
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng et al.ACL 2026 · 30 citations
- Annotator-Centric Active Learning for Subjective NLP TasksMichiel van der Meer, Neele Falk, Pradeep K. Murukannaiah, Enrico LiscioEMNLP 2024 · 3 citations
- Confidence-based Ensembling of Perspective-aware ModelsSilvia Casola, Soda Marem Lo, Valerio Basile, Simona Frenda et al.EMNLP 2023 · 2 citations
- Learning Personalized Alignment for Evaluating Open-ended Text GenerationDanqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang et al.EMNLP 2024 · 2 citations
