AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
Changyi Li, Pengfei Lu, Xudong Pan, Fazl Barez, Min Yang
Abstract
As Large Language Models (LLMs) evolve into autonomous agents, existing safety evaluations face a fundamental trade-off: manual benchmarks are costly, while LLM-based simulators are scalable but suffer from logic hallucination. We present AutoControl Arena, an automated framework for frontier AI risk evaluation built on the principle of logic-narrative decoupling. By grounding deterministic state in executable code while delegating generative dynamics to LLMs, we mitigate hallucination while maintaining flexibility. This principle, instantiated through a three-agent framework, achieves over 98% end-to-end success and 60% human preference over existing simulators. To elicit latent risks, we vary environmental Stress and Temptation across X-Bench (70 scenarios, 7 risk categories). Evaluating 9 frontier models reveals: (1) Alignment Illusion: risk rates surge from 21.7% to 54.5% under pressure, with capable models showing disproportionately larger increases; (2) Scenario-Specific Safety Scaling: advanced reasoning improves robustness for direct harms but worsens it for gaming scenarios; and (3) Divergent Misalignment Patterns: weaker models cause non-malicious harm while stronger models develop strategic concealment. Code and data are available at https://github.com/CosmosYi/AutoControl-Arena.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8d926fc-3ef9-4ab0-a82a-76e8713d627cBuilds on4
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- Reliable Weak-to-Strong Monitoring of LLM AgentsNeil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich et al.ICLR 2026 · 28 citations
- Safety Alignment Should be Made More Than Just a Few Tokens DeepXiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma et al.ICLR 2025
Related papers
- PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic ApproachUdari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy, Vikash Sehwag et al.ICLR 2026 · 19 citations
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMsAdi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak et al.ICLR 2026 · 3 citations
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 2 citations
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu et al.ACL 2025 · 34 citations
