Evaluating Contextual Illegality: AI Compliance in Corporate Law Scenarios
Hilal Aka, Joe Kwon, Noam Kolt
摘要
While AI models often refuse explicitly unlawful requests, in real-world scenarios illegality often depends on context. We evaluate frontier models on contextual illegality across four corporate law scenarios in which routine actions—editing documents, trading stock, requesting payment, approving communications—become unlawful due to circumstances such as pending investigations or bankruptcy filings. We study both chat and agentic settings and compare results to a human baseline. The best-performing models consistently followed lawful requests and refused unlawful requests, though performance varied substantially between different scenarios and models. We also identify distinct failure modes, such as excessive refusal of lawful requests, and find higher performance in reasoning models and agentic environments. By studying contextual illegality in these controlled environments, we develop a methodology that can be extended to evaluate the legal compliance of AI models in additional scenarios and domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsMaksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas 等ICLR 2025 · 被引用 4 次
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language ModelsGuangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland 等ICML 2025
相关 Paper
- Implicit Intelligence - Evaluating Agents on What Users Don’t SayVed Sirdeshmukh, Marc WetterICML 2026
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren 等ICML 2026 · 被引用 5 次
- ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World HarmKefan Song, Yanjun QiICML 2026
