Value Profiles for Encoding Human Variation
Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah D. Goodman, Verena Rieser
Abstract
Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using natural language value profiles -descriptions of underlying values compressed from in-context demonstrations -along with a steerable decoder model that estimates individual ratings from a rater representation. To measure the predictive information in a rater representation, we introduce an information-theoretic methodology and find that demonstrations contain the most information, followed by value profiles, then demographics. However, value profiles effectively compress the useful information from demonstrations (>70% information preservation) and offer advantages in terms of scrutability, interpretability, and steerability. Furthermore, clustering value profiles to identify similarly behaving individuals better explains rater variation than the most predictive demographic groupings. Going beyond test set performance, we show that the decoder predictions change in line with semantic profile differences, are wellcalibrated, and can help explain instance-level disagreement by simulating an annotator population. These results demonstrate that value profiles offer novel, predictive ways to describe individual variation beyond demographics or group information. Is high immigration good for the UK? Should we ban cars in city centers? Age: 40 Gender: Female … Income: £70k Demographics Age: 40 Gender: Female … Income: £70k Demographics Should we ban cars in city centers? No Yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb3514e7-2011-4bb1-989b-a34486cb17b3Cited by top-tier papers7
- Spectrum Tuning: Post-Training for Distributional Coverage and In-Context SteerabilityTaylor Sorensen, Benjamin Newman, Jared Moore, Chan Young Park et al.ICLR 2026 · 19 citations
- Pairwise Calibrated Rewards for Pluralistic AlignmentDaniel Halpern, Evi Micha, Ariel D. Procaccia, Itai ShapiraNeurIPS 2025 · 15 citations
- PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI HarmJingjing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin et al.ICLR 2026 · 6 citations
- Improving the Distributional Alignment of LLMs using SupervisionGauri Kambhatla, Sanjana Gautam, Angela Zhang, Alexander Liu et al.ACL 2026 · 1 citation
- SOLAR: Towards Characterizing Subjectivity of Individuals through Modeling Value Conflicts and Trade-offsYounghun Lee, Dan GoldwasserEMNLP 2025
Builds on18
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
Related papers
- Do Differences in Values Influence Disagreements in Online Discussions?Michiel van der Meer, Piek Vossen, Catholijn M. Jonker, Pradeep K. MurukannaiahEMNLP 2023 · 2 citations
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- Can Language Models Reason about Individualistic Human Values and Preferences?Liwei Jiang, Taylor Sorensen, Sydney Levine, Yejin ChoiACL 2025
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language ModelsHanze Guo, Jing Yao, Xiao Zhou, Xiaoyuan Yi et al.NeurIPS 2025 · 6 citations
- Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive ReasoningHaijiang Liu, Qiyuan Li, Chao Gao, Yong Cao et al.EMNLP 2025
