Prompting is not a substitute for probability measurements in large language models
Jennifer Hu, Roger Levy
摘要
Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgment. In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models' linguistic knowledge. Broadly, we find that LLMs' metalinguistic judgments are inferior to quantities directly derived from representations. Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities. Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization. Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Bias in Language Models: Beyond Trick Tests and Towards RUTEd EvaluationKristian Lum, Jacy Reese Anthis, Kevin Robinson, Chirag Nagpal 等ACL 2025 · 被引用 41 次
- Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory informationMurathan Kurfali, Robert ÖstlingEMNLP 2025 · 被引用 4 次
- Paraphrase Types Elicit Prompt Engineering CapabilitiesJan Philip Wahle, Terry Ruas, Yang Xu, Bela GippEMNLP 2024 · 被引用 4 次
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 被引用 3 次
- Cognitive models can reveal interpretable value trade-offs in language modelsSonia Krishna Murthy, Rosie Zhao, Jennifer Hu, Sham M. Kakade 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
相关 Paper
- Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying ProbesSharan Maiya, Yinhong Liu, Ramit Debnath, Anna KorhonenACL 2025 · 被引用 4 次
- Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' MemoryXingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang 等ICLR 2026 · 被引用 1 次
- Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model EvaluationXunjian Yin, Xu Zhang, Jie Ruan, Xiaojun WanACL 2024
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko 等ACL 2023 · 被引用 45 次
- Shared Lexical Task Representations Explain Behavioral Variability In LLMsZhuonan Yang, Jacob Xiaochen Li, Francisco Velez, Eric Todd 等ICML 2026
