Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
Nora Petrova, Andrew Gordon, Enzo Blindow
Abstract
The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from unrepresentative sampling, superficial assessment depth, and single-metric reductionism. To address these issues, we introduce HUMAINE, a framework for multidimensional, demographically aware measurement of human-AI interaction. We collected multi-turn, naturalistic conversations from 23,404 participants that were stratified across 22 demographic groups, both in the US and UK, to evaluate 28 state-of-the-art models across five human-centric dimensions. We use a hierarchical Bayesian Bradley-Terry-Davidson (BTD) model, with post-stratification to census data, and our analysis reveals three key insights. (1) We establish a clear performance hierarchy where google/gemini-2.5-pro ranks first overall, with a 95.6% posterior probability of being the top-ranked model. (2) We uncover significant preference heterogeneity, with user age emerging as the primary demographic axis of disagreement; a model's perceived rank can shift substantially across age groups, exposing failures in generalisation that unrepresentative samples typically mask. (3) We quantify the vast difference in discriminative power across evaluation dimensions, with ambiguous qualities like Trust, Ethics &Safety showing a 65% tie rate, in stark contrast to the decisive 10% tie rate for Overall Winner. Our work emphasises the need for a more multidimensional, demographically aware perspective in LLM evaluation. We release our complete dataset, interactive leaderboard, and open-source framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 96 citations
- Dissecting Human and LLM PreferencesJunlong Li, Fan Zhou, Shichao Sun, Yikai Zhang et al.ACL 2024 · 1 citation
Related papers
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language ModelsLujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi et al.ICLR 2026 · 40 citations
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller et al.ACL 2025
- Talk2Code: A Multi-Turn Interaction Benchmark with Dual-Track Evaluation for Code GenerationWeibin Yang, Liangru Xie, Jieyun Cai, Yuxiang Yan et al.AAAI 2026
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
- To Mask or to Mirror: Human-AI Alignment in Collective ReasoningCrystal Qian, Aaron T. Parisi, Clémentine Bouleau, Vivian Tsai et al.EMNLP 2025 · 1 citation
