CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such
Abstract
Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refusal of individual problematic queries, which overlooks the importance of the context where the query occurs and may cause undesired refusal of queries under safe contexts that diminish user experience. Addressing this gap, we introduce CASE-Bench, a Context-Aware SafEty Benchmark that integrates context into safety assessments of LLMs. CASE-Bench assigns distinct, formally described contexts to categorized queries based on Contextual Integrity theory. Additionally, in contrast to previous studies which mainly rely on majority voting from just a few annotators, we recruited a sufficient number of annotators necessary to ensure the detection of statistically significant differences among the experimental conditions based on power analysis. Our extensive analysis using CASE-Bench on various open-source and commercial LLMs reveals a substantial and significant influence of context on human judgments (p <0.0001 from a z-test), underscoring the necessity of context in safety evaluations. We also identify notable mismatches between human judgments and LLM responses, particularly in commercial models within safe contexts. 1 Code and data used in the paper are available at https://github.com/BriansIDP/CASEBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eaba11e8-bef9-4826-ad43-41c4bd214f2bCited by top-tier papers1
Ask how each one uses itBuilds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li et al.ICLR 2024 · 481 citations
Related papers
- PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal ComplianceHaoran Li, Wenbin Hu, Huihao Jing, Yulin Chen et al.ACL 2025
- LongSafety: Evaluating Long-Context Safety of Large Language ModelsYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui et al.ACL 2025 · 6 citations
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas et al.ICLR 2025
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz et al.ACL 2025 · 17 citations
- SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalTinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang et al.ICLR 2025
