Shh, don't say that! Domain Certification in LLMs
Cornelius Emde, Alasdair Paren, Preetham Arvind, Maxime Guillaume Kayser, Tom Rainforth, Thomas Lukasiewicz, Philip Torr, Adel Bibi
摘要
Large language models (LLMs) are often deployed to perform constrained tasks, with narrow domains. For example, customer support bots can be built on top of LLMs, relying on their broad language understanding and capabilities to enhance performance. However, these LLMs are adversarially susceptible, potentially generating outputs outside the intended domain. To formalize, assess, and mitigate this risk, we introduce domain certification; a guarantee that accurately characterizes the out-of-domain behavior of language models. We then propose a simple yet effective approach, which we call VALID that provides adversarial bounds as a certificate. Finally, we evaluate our method across a diverse set of datasets, demonstrating that it yields meaningful certificates, which bound the probability of out-of-domain samples tightly with minimum penalty to refusal behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Lifelong Safety Alignment for Language ModelsHaoyu Wang, Yifei Zhao, Zeyu Qin, Chao Du 等NeurIPS 2025 · 被引用 18 次
- MIP against Agent: Malicious Image Patches Hijacking Multimodal OS AgentsLukas Aichberger, Alasdair Paren, Guohao Li, Philip H. S. Torr 等NeurIPS 2025 · 被引用 12 次
- Safety Depth in Large Language Models: A Markov Chain PerspectiveChing-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song ChenNeurIPS 2025 · 被引用 2 次
- How Catastrophic is Your LLM? Certifying Risks in ConversationChengxiao Wang, Isha Chaudhary, Qian Hu, Weitong Ruan 等ICLR 2026 · 被引用 1 次
- Beyond the Known: An Unknown-Aware Large Language Model for Open-Set Text ClassificationXi Chen, Chuan Qin, Ziqi Wang, Shasha Hu 等ICLR 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi 等ICLR 2025 · 被引用 3 次
- Learning Safety Constraints for Large Language ModelsXin Chen, Yarden As, Andreas KrauseICML 2025
- Data to Defense: The Role of Curation in Aligning Large Language Models Against Safety CompromiseXiaoqun Liu, Jiacheng Liang, Luoxi Tang, Muchao Ye 等EMNLP 2025
- Semantic Robustness Certification for Vision-Language ModelsPeiyu Yang, Paul MONTAGUE, Feng Liu, Andrew C. Cullen 等ICML 2026
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao 等ICML 2026 · 被引用 6 次
