Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
Xiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu, Kuntai Cai, Yan Xiao, Jin Song Dong
摘要
The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we introduce a novel framework POLARIS that brings the rigor of specification-based software testing to AI safety. POLARIS first compiles unstructured natural-language policies into First-Order Logic (FOL) representations, establishing a traceable link between high-level rules and concrete test cases. This formalization enables the construction of a Semantic Policy Graph, where complex policy violation scenarios are encoded as traversable paths. By systematically exploring this graph, POLARIS uncovers compositional violation patterns, which are then instantiated into executable natural-language test queries, enabling coverage-driven and reproducible safety testing. Experiments demonstrate that POLARIS achieves higher policy coverage and attack success counts compared to established baselines. Crucially, by bridging formal methods and AI safety, POLARIS provides a principled, automated approach to ensuring LLMs adhere to safety-critical policies with verifiable traceability. We release our code at https://github.com/huac-lxy/POLARIS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Talk2Care: An LLM-based Voice Assistant for Communication between Healthcare Providers and Older AdultsZiqi Yang, Xuhai Xu, Bingsheng Yao, Ethan Rogers 等UbiComp 2024 · 被引用 116 次
- Curiosity-driven Red-teaming for Large Language ModelsZhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang 等ICLR 2024 · 被引用 84 次
- Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual UnderstandingHaneul Yoo, Yongjin Yang, Hwaran LeeACL 2025 · 被引用 27 次
- Property-Based Testing in PracticeHarrison Goldstein, Joseph W. Cutler, Daniel Dickstein, Benjamin C. Pierce 等ICSE 2024 · 被引用 21 次
相关 Paper
- CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional SafetyLu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An 等ISSTA 2026
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack IntegrationAndy Zhou, Kevin Wu, Francesco Pinto, Zhaorun Chen 等NeurIPS 2025 · 被引用 46 次
- Refusal-Aware Red Teaming: Exposing Inconsistency in Safety EvaluationsYongkang Chen, Xiaohu Du, Xiaotian Zou, Chongyang Zhao 等EMNLP 2025 · 被引用 2 次
- Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language ModelsKai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang 等ACL 2026
- Red Teaming LLMs via Linguistic-Aware FuzzingShuai Yuan, Nian Luo, Jingling Sun, Yihao Huang 等FSE 2026
