Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
Paulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der Schaar
摘要
The predominant de facto paradigm of testing ML models relies on either using only held-out data to compute aggregate evaluation metrics or by assessing the performance on different subgroups. However, such data-only testing methods operate under the restrictive assumption that the available empirical data is the sole input for testing ML models, disregarding valuable contextual information that could guide model testing. In this paper, we challenge the go-to approach of data-only testing and introduce context-aware testing (CAT) which uses context as an inductive bias to guide the search for meaningful model failures. We instantiate the first CAT system, SMART Testing, which employs large language models to hypothesize relevant and likely failures, which are evaluated on data using a self-falsification mechanism. Through empirical evaluations in diverse settings, we show that SMART automatically identifies more relevant and impactful failures than alternatives, demonstrating the potential of CAT as a testing paradigm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language ModelsPaulius Rauba, Mihaela van der SchaarICLR 2026 · 被引用 3 次
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang 等ICML 2025
- No More, No Less: Least-Privilege Language ModelsPaulius Rauba, Dominykas Seputis, Patrikas Vanagas, Mihaela van der SchaarICML 2026
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
它引用的顶会 Paper12
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- Model Patching: Closing the Subgroup Performance Gap with Data AugmentationKaran Goel, Albert Gu, Sharon Li, Christopher RéICLR 2021 · 被引用 131 次
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar 等ICLR 2024 · 被引用 114 次
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 被引用 61 次
相关 Paper
- Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang 等ISSTA 2026
- CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional SafetyLu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An 等ISSTA 2026
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak 等NeurIPS 2025 · 被引用 11 次
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang 等ICML 2026 · 被引用 4 次
- Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph ReasoningMuzhi Li, Cehao Yang, Chengjin Xu, Zixing Song 等AAAI 2025 · 被引用 7 次
