Measuring Intent Comprehension in LLMs
Nadav Kunievsky, James Evans
Abstract
People judge interactions with large language models (LLMs) as successful when outputs match what they want, not what they type. Yet LLMs are trained to predict the next token solely from text input, not underlying intent. Because written language is an imperfect proxy for intent, and correlations between phrasing and desired outcomes can break down in training data, models that rely too heavily on surface cues may respond inconsistently to semantically equivalent prompts. This makes it essential to evaluate whether LLMs can reliably infer user intent-especially in highstakes settings where robustness and generalization are critical. We introduce a formal framework for assessing intent comprehension in LLMs: whether a model demonstrates robust understanding of user intent by producing consistent outputs across semantically equivalent prompts while differentiating between prompts with distinct intents. Our evaluation approach is based on a variance decomposition of model responses into three components: variability due to user intent, user articulation, and the residual uncertainty. Models that understand what users want, and are not overly sensitive to textual cues, should attribute most output variance to intent differences, rather than articulation style. Applying this framework across diverse five domains, we find that, within the five LLaMA and Gemma models we evaluate, larger models typically assign a greater share of variance to intent, indicating stronger comprehension of intent, although gains are uneven and often modest with increasing model size. These results motivate moving beyond accuracy-only benchmarks toward semantic diagnostics that directly assess whether models understand what users intend.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3d5e142-3161-4186-8904-b7b9f8fc9b09Cited by top-tier papers2
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
- Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language ModelsZhengshuyuan Tian, Wanling Gao, Chuanxin Lan, Chenxi Wang et al.ICML 2026
Builds on1
Related papers
- On the Worst Prompt Performance of Large Language ModelsBowen Cao, Deng Cai, Zhisong Zhang, Yuexian Zou et al.NeurIPS 2024 · 65 citations
- DiscoverLLM: From Executing Intents to Discovering ThemTae Soo Kim, Yoonjoo Lee, Jaesang Yu, John Chung et al.ICML 2026
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing StylesKimberly Le Truong, Riccardo Fogliato, Hoda Heidari, Steven WuEMNLP 2025 · 2 citations
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning et al.ICML 2025
- Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal ResponsesSugyeong Eo, Heuiseok LimACL 2026
