PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
Jingjing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin, Anne Collins, Maarten Sap, Sydney Levine
Abstract
Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PLURIHARMS, a benchmark designed to systematically study human harm judgments across two key dimensions-the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PLURIHARMS, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith et al.ICLR 2026 · 41 citations
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent ApproachYuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian et al.NeurIPS 2025 · 20 citations
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 8 citations
- Value Profiles for Encoding Human VariationTaylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler et al.EMNLP 2025 · 2 citations
Related papers
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.ACL 2025
- VITAL: A New Dataset for Benchmarking Pluralistic Alignment in HealthcareAnudeex Shetty, Amin Beheshti, Mark Dras, Usman NaseemACL 2025 · 13 citations
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMsAdi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak et al.ICLR 2026 · 3 citations
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated TeachersYilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin et al.AAAI 2026 · 1 citation
- Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust EvaluationHuije Lee, Jisu Shin, Hoyun Song, Changgeon Ko et al.ACL 2026
