Shh, don't say that! Domain Certification in LLMs
Cornelius Emde, Alasdair Paren, Preetham Arvind, Maxime Guillaume Kayser, Tom Rainforth, Thomas Lukasiewicz, Philip Torr, Adel Bibi
Abstract
Large language models (LLMs) are often deployed to perform constrained tasks, with narrow domains. For example, customer support bots can be built on top of LLMs, relying on their broad language understanding and capabilities to enhance performance. However, these LLMs are adversarially susceptible, potentially generating outputs outside the intended domain. To formalize, assess, and mitigate this risk, we introduce domain certification; a guarantee that accurately characterizes the out-of-domain behavior of language models. We then propose a simple yet effective approach, which we call VALID that provides adversarial bounds as a certificate. Finally, we evaluate our method across a diverse set of datasets, demonstrating that it yields meaningful certificates, which bound the probability of out-of-domain samples tightly with minimum penalty to refusal behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6f9d0de-826d-41a4-b003-5e370f30bca9Cited by top-tier papers5
- Lifelong Safety Alignment for Language ModelsHaoyu Wang, Yifei Zhao, Zeyu Qin, Chao Du et al.NeurIPS 2025 · 18 citations
- MIP against Agent: Malicious Image Patches Hijacking Multimodal OS AgentsLukas Aichberger, Alasdair Paren, Guohao Li, Philip H. S. Torr et al.NeurIPS 2025 · 12 citations
- Safety Depth in Large Language Models: A Markov Chain PerspectiveChing-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song ChenNeurIPS 2025 · 2 citations
- How Catastrophic is Your LLM? Certifying Risks in ConversationChengxiao Wang, Isha Chaudhary, Qian Hu, Weitong Ruan et al.ICLR 2026 · 1 citation
- Beyond the Known: An Unknown-Aware Large Language Model for Open-Set Text ClassificationXi Chen, Chuan Qin, Ziqi Wang, Shasha Hu et al.ICLR 2026
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi et al.ICLR 2025 · 3 citations
- Learning Safety Constraints for Large Language ModelsXin Chen, Yarden As, Andreas KrauseICML 2025
- Data to Defense: The Role of Curation in Aligning Large Language Models Against Safety CompromiseXiaoqun Liu, Jiacheng Liang, Luoxi Tang, Muchao Ye et al.EMNLP 2025
- Semantic Robustness Certification for Vision-Language ModelsPeiyu Yang, Paul MONTAGUE, Feng Liu, Andrew C. Cullen et al.ICML 2026
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao et al.ICML 2026 · 6 citations
