Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Siqi Shen, Mehar Singh, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Rada Mihalcea
摘要
The value orientation of Large Language Models (LLMs) has been extensively studied, as it can shape user experiences across demographic groups. However, two key challenges remain: (1) the lack of systematic comparison across value probing strategies, despite the Multiple Choice Question (MCQ) setting being vulnerable to perturbations, and (2) the uncertainty over whether probed values capture in-context information or predict models' real-world actions. In this paper, we systematically compare three widely used value probing methods: token likelihood, sequence perplexity, and text generation. Our results show that all three methods exhibit large variances under non-semantic perturbations in prompts and option formats, with sequence perplexity being the most robust overall. We further introduce two tasks to assess expressiveness: demographic prompting, testing whether probed values adapt to cultural context; and value-action agreement, testing the alignment of probed values with value-based actions. We find that demographic context has little effect on the text generation method, and probed values only weakly correlate with action preferences across all methods. Our work highlights the instability and the limited expressive power of current value probing methods, calling for more reliable LLM value representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 被引用 27 次
- Measuring Human and AI Values Based on Generative Psychometrics with Large Language ModelsHaoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang 等AAAI 2025 · 被引用 18 次
相关 Paper
- Do LLMs have Consistent Values?Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson 等ICLR 2025
- Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic PromptingSagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali 等EMNLP 2024 · 被引用 3 次
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value DifferenceJing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu 等ICLR 2026 · 被引用 4 次
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 被引用 116 次
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 被引用 5 次
