Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
Paulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der Schaar
Abstract
The predominant de facto paradigm of testing ML models relies on either using only held-out data to compute aggregate evaluation metrics or by assessing the performance on different subgroups. However, such data-only testing methods operate under the restrictive assumption that the available empirical data is the sole input for testing ML models, disregarding valuable contextual information that could guide model testing. In this paper, we challenge the go-to approach of data-only testing and introduce context-aware testing (CAT) which uses context as an inductive bias to guide the search for meaningful model failures. We instantiate the first CAT system, SMART Testing, which employs large language models to hypothesize relevant and likely failures, which are evaluated on data using a self-falsification mechanism. Through empirical evaluations in diverse settings, we show that SMART automatically identifies more relevant and impactful failures than alternatives, demonstrating the potential of CAT as a testing paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f90515ec-ab80-47dc-8883-fa7850fbee12Cited by top-tier papers6
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language ModelsPaulius Rauba, Mihaela van der SchaarICLR 2026 · 3 citations
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang et al.ICML 2025
- No More, No Less: Least-Privilege Language ModelsPaulius Rauba, Dominykas Seputis, Patrikas Vanagas, Mihaela van der SchaarICML 2026
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
Builds on12
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Model Patching: Closing the Subgroup Performance Gap with Data AugmentationKaran Goel, Albert Gu, Sharon Li, Christopher RéICLR 2021 · 131 citations
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar et al.ICLR 2024 · 114 citations
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 61 citations
Related papers
- Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang et al.ISSTA 2026
- CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional SafetyLu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An et al.ISSTA 2026
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak et al.NeurIPS 2025 · 11 citations
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang et al.ICML 2026 · 4 citations
- Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph ReasoningMuzhi Li, Cehao Yang, Chengjin Xu, Zixing Song et al.AAAI 2025 · 7 citations
