Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering
Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee
Abstract
As Large Language Models (LLMs) are increasingly deployed in sensitive domains such as enterprise and government, ensuring that they adhere to user-defined security policies within context is critical-especially with respect to information non-disclosure. While prior LLM studies have focused on general safety and socially sensitive data, large-scale benchmarks for contextual security preservation against attacks remain lacking. To address this, we introduce a novel large-scale benchmark dataset, CoPriva, evaluating LLM adherence to contextual non-disclosure policies in question answering. Derived from realistic contexts, our dataset includes explicit policies and queries designed as direct and challenging indirect attacks seeking prohibited information. We evaluate 10 LLMs on our benchmark and reveal a significant vulnerability: many models violate user-defined policies and leak sensitive information. This failure is particularly severe against indirect attacks, highlighting a critical gap in current LLM safety alignment for sensitive applications. Our analysis reveals that while models can often identify the correct answer to a query, they struggle to incorporate policy constraints during generation. In contrast, they exhibit a partial ability to revise outputs when explicitly prompted. Our findings underscore the urgent need for more robust methods to guarantee contextual security. 1 * Equal contribution. † Corresponding author. 1 https://github.com/hwanchang00/CoPri va Do not disclose speech recognition feature debate. (…) Industrial Designer: If we aim for the younger people , and there will be a lot of features like LCD or the speech recognising , the cost will be higher. I think we don't have that in our budget. Project Manager: I think the LCD is cheaper than speech recognition. So I think that can be a good option. LCD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 102f4ece-e17a-46ef-93d7-273d140072bbBuilds on5
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin et al.EMNLP 2024 · 38 citations
- GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity TheoryWei Fan, Haoran Li, Zheye Deng, Weiqi Wang et al.EMNLP 2024 · 7 citations
Related papers
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
- Can We Infer Confidential Properties of Training Data from LLMs?Pengrun Huang, Chhavi Yadav, Kamalika Chaudhuri, Ruihan WuNeurIPS 2025 · 7 citations
- A Benchmark for Semantic Sensitive Information in LLMs OutputsQingjie Zhang, Han Qiu, Di Wang, Yiming Li et al.ICLR 2025
- Security Attacks on LLM-based Code Completion ToolsWen Cheng, Ke Sun, Xinyu Zhang, Wei WangAAAI 2025 · 21 citations
- CIMemories: A Compositional Benchmark For Contextual Integrity In LLMsNiloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov et al.ICLR 2026 · 10 citations
