Implicit Intelligence - Evaluating Agents on What Users Don’t Say
Ved Sirdeshmukh, Marc Wetter
Abstract
Real-world requests to AI agents are fundamentally underspecified. Natural human communication relies on shared context and unstated constraints that speakers expect listeners to infer. Current agentic benchmarks test explicit instruction-following but fail to evaluate whether agents can reason about implicit requirements spanning accessibility needs, privacy boundaries, catastrophic risks, and contextual constraints. We present Implicit Intelligence , an evaluation framework testing whether AI agents can move beyond prompt-following to become genuine goal-fulfillers, paired with Agent-as-a-World (AaW) , a harness where interactive worlds are defined in human-readable YAML files and simulated by language models. Our scenarios feature apparent simplicity in user requests, hidden complexity in correct solutions, and discoverability of constraints through environmental exploration. Evaluating 16 frontier and open-weight models across 205 scenarios, we find that even the best-performing model achieves only 48.3% scenario pass rate, revealing substantial room for improvement in bridging the gap between literal instruction-following and human-like contextual reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e08478c-ff70-4b2b-814f-df7664f1570bBuilds on5
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language ModelsTao Zou, Xinghua Zhang, Haiyang Yu, Minzheng Wang et al.EMNLP 2025 · 1 citation
Related papers
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- Benchmarking World-Model Learning with Environment-Level QueriesArchana Warrier, Dat Nguyen, Michelangelo Naim, Moksh Jain et al.ICML 2026 · 3 citations
- Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsQianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao et al.EMNLP 2025
- Evaluating Contextual Illegality: AI Compliance in Corporate Law ScenariosHilal Aka, Joe Kwon, Noam KoltICML 2026
- InteractComp: Evaluating Search Agents With Ambiguous QueriesMingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong et al.ICML 2026 · 11 citations
