Prompting is not a substitute for probability measurements in large language models
Jennifer Hu, Roger Levy
Abstract
Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgment. In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models' linguistic knowledge. Broadly, we find that LLMs' metalinguistic judgments are inferior to quantities directly derived from representations. Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities. Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization. Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1751d6af-9cc3-4d9c-9f40-511a34d7dc34Cited by top-tier papers16
- Bias in Language Models: Beyond Trick Tests and Towards RUTEd EvaluationKristian Lum, Jacy Reese Anthis, Kevin Robinson, Chirag Nagpal et al.ACL 2025 · 41 citations
- Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory informationMurathan Kurfali, Robert ÖstlingEMNLP 2025 · 4 citations
- Paraphrase Types Elicit Prompt Engineering CapabilitiesJan Philip Wahle, Terry Ruas, Yang Xu, Bela GippEMNLP 2024 · 4 citations
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 3 citations
- Cognitive models can reveal interpretable value trade-offs in language modelsSonia Krishna Murthy, Rosie Zhao, Jennifer Hu, Sham M. Kakade et al.ICLR 2026 · 2 citations
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
Related papers
- Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying ProbesSharan Maiya, Yinhong Liu, Ramit Debnath, Anna KorhonenACL 2025 · 4 citations
- Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' MemoryXingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang et al.ICLR 2026 · 1 citation
- Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model EvaluationXunjian Yin, Xu Zhang, Jie Ruan, Xiaojun WanACL 2024
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko et al.ACL 2023 · 45 citations
- Shared Lexical Task Representations Explain Behavioral Variability In LLMsZhuonan Yang, Jacob Xiaochen Li, Francisco Velez, Eric Todd et al.ICML 2026
