Statistical Hypothesis Testing for Auditing Robustness in Language Models
Paulius Rauba, Qiyao Wei, Mihaela van der Schaar
Abstract
Consider the problem of testing whether the outputs of a large language model (LLM) system change under an arbitrary intervention, such as an input perturbation or changing the model variant. We cannot simply compare two LLM outputs since they might differ due to the stochastic nature of the system, nor can we compare the entire output distribution due to computational intractability. While existing methods for analyzing text-based outputs exist, they focus on fundamentally different problems, such as measuring bias or fairness. To this end, we introduce distribution-based perturbation analysis, a framework that reformulates LLM perturbation analysis as a frequentist hypothesis testing problem. We construct empirical null and alternative output distributions within a low-dimensional semantic similarity space via Monte Carlo sampling, enabling tractable inference without restrictive distributional assumptions. The framework is (i) modelagnostic, (ii) supports the evaluation of arbitrary input perturbations on any black-box LLM, (iii) yields interpretable p-values; (iv) supports multiple perturbations via controlled error rates; and (v) provides scalar effect sizes. We demonstrate the usefulness of the framework across multiple case studies, showing how we can quantify response changes, measure true/false positive rates, and evaluate alignment with reference models. Above all, we see this as a reliable frequentist hypothesis testing framework for LLM auditing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language ModelsPaulius Rauba, Mihaela van der SchaarICLR 2026 · 3 citations
- No More, No Less: Least-Privilege Language ModelsPaulius Rauba, Dominykas Seputis, Patrikas Vanagas, Mihaela van der SchaarICML 2026
Builds on2
Related papers
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
- Model Equality Testing: Which Model is this API Serving?Irena Gao, Percy Liang, Carlos GuestrinICLR 2025
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestXiaoyuan Zhu, Yaowen Ye, Tianyi Qiu, Hanlin Zhu et al.ICLR 2026 · 22 citations
- Enhancing Hallucination Detection through Noise InjectionLitian Liu, Reza Pourreza, Sunny Panchal, Apratim Bhattacharyya et al.ICLR 2026 · 19 citations
- Harnessing Non-Adversarial Robustness in Large Language ModelsQinghua Zhou, Ellina Aleshina, Andrey Lovyagin, Oleg Somov et al.ICML 2026
