Value Profiles for Encoding Human Variation
Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah D. Goodman, Verena Rieser
摘要
Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using natural language value profiles -descriptions of underlying values compressed from in-context demonstrations -along with a steerable decoder model that estimates individual ratings from a rater representation. To measure the predictive information in a rater representation, we introduce an information-theoretic methodology and find that demonstrations contain the most information, followed by value profiles, then demographics. However, value profiles effectively compress the useful information from demonstrations (>70% information preservation) and offer advantages in terms of scrutability, interpretability, and steerability. Furthermore, clustering value profiles to identify similarly behaving individuals better explains rater variation than the most predictive demographic groupings. Going beyond test set performance, we show that the decoder predictions change in line with semantic profile differences, are wellcalibrated, and can help explain instance-level disagreement by simulating an annotator population. These results demonstrate that value profiles offer novel, predictive ways to describe individual variation beyond demographics or group information. Is high immigration good for the UK? Should we ban cars in city centers? Age: 40 Gender: Female … Income: £70k Demographics Age: 40 Gender: Female … Income: £70k Demographics Should we ban cars in city centers? No Yes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Spectrum Tuning: Post-Training for Distributional Coverage and In-Context SteerabilityTaylor Sorensen, Benjamin Newman, Jared Moore, Chan Young Park 等ICLR 2026 · 被引用 19 次
- Pairwise Calibrated Rewards for Pluralistic AlignmentDaniel Halpern, Evi Micha, Ariel D. Procaccia, Itai ShapiraNeurIPS 2025 · 被引用 15 次
- PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI HarmJingjing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin 等ICLR 2026 · 被引用 6 次
- Improving the Distributional Alignment of LLMs using SupervisionGauri Kambhatla, Sanjana Gautam, Angela Zhang, Alexander Liu 等ACL 2026 · 被引用 1 次
- SOLAR: Towards Characterizing Subjectivity of Individuals through Modeling Value Conflicts and Trade-offsYounghun Lee, Dan GoldwasserEMNLP 2025
它引用的顶会 Paper18
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart 等ICLR 2020 · 被引用 211 次
相关 Paper
- Do Differences in Values Influence Disagreements in Online Discussions?Michiel van der Meer, Piek Vossen, Catholijn M. Jonker, Pradeep K. MurukannaiahEMNLP 2023 · 被引用 2 次
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 被引用 11 次
- Can Language Models Reason about Individualistic Human Values and Preferences?Liwei Jiang, Taylor Sorensen, Sydney Levine, Yejin ChoiACL 2025
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language ModelsHanze Guo, Jing Yao, Xiao Zhou, Xiaoyuan Yi 等NeurIPS 2025 · 被引用 6 次
- Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive ReasoningHaijiang Liu, Qiyuan Li, Chao Gao, Yong Cao 等EMNLP 2025
