CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety
Lu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An, Guangyu Shen, Zhou Xuan, Xuan Chen, Xiangyu Zhang
摘要
Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows where individual requests are composed into complex behavior. This paper introduces compositional safety, the property that an LLM remains safe not only against isolated malicious prompts, but also under structured, long-horizon decompositions of harmful intents. We propose CAST, a systematic testing framework designed to evaluate the compositional safety of LLMs in the domain of malicious code. Drawing inspiration from modern compiler infrastructures, CAST decouples test case generation from test execution using a novel intermediate representation, CAIR. This architecture allows the framework to automatically refine high-level testing intents into granular sub-tasks that serve as unit tests for the model’s alignment. These components are subsequently instantiated by the SUT and reassembled according to the CAIR control structure. The resulting artifact is then evaluated by intent-fulfillment scoring and, for the severity subset, external behavioral detectors and manual inspection. We evaluate CAST on four state-of-the-art LLMs across three security-critical testbeds. Our results demonstrate that CAST systematically exposes severe safety violations in strongly aligned models that resist conventional red-teaming, achieving up to a 365% increase in successful test cases compared to baseline testing strategies
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Inverting the Shield: Systematically Generating Safety Tests from Policy SpecificationsXiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu 等ACL 2026
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-BreakingYifan Huang, Xiaojun Jia, Wenbo Guo, Yuqiang Sun 等FSE 2026
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak AttacksMingqian Feng, Xiaodong Liu, Weiwei Yang, Jialin Song 等ICLR 2026 · 被引用 13 次
- LongSafety: Evaluating Long-Context Safety of Large Language ModelsYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui 等ACL 2025 · 被引用 6 次
- How Catastrophic is Your LLM? Certifying Risks in ConversationChengxiao Wang, Isha Chaudhary, Qian Hu, Weitong Ruan 等ICLR 2026 · 被引用 1 次
