AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li
Abstract
Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, existing public benchmarks often define safety categories based on previous literature, intuitions, or common sense, leading to disjointed sets of categories for risks specified in recent regulations and policies, which makes it challenging to evaluate and compare FMs across these benchmarks. To bridge this gap, we introduce AIR-Bench 2024, the first AI safety benchmark aligned with emerging government regulations and company policies, following the regulation-based safety categories grounded in our AI risks study, AIR 2024. AIR 2024 decomposes 8 government regulations and 16 company policies into a four-tiered safety taxonomy with 314 granular risk categories in the lowest tier. AIR-Bench 2024 contains 5,694 diverse prompts spanning these categories, with manual curation and human auditing to ensure quality. We evaluate leading language models on AIR-Bench 2024, uncovering insights into their alignment with specified safety concerns. By bridging the gap between public benchmarks and practical AI risks, AIR-Bench 2024 provides a foundation for assessing model safety across jurisdictions, fostering the development of safer and more responsible AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2781e2b1-dc0a-4535-9e2a-a6a26620bc8bCited by top-tier papers6
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS EnvironmentsZeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning et al.ICLR 2026 · 46 citations
- Towards Policy-Adaptive Image Guardrail: Benchmark and MethodCaiyong Piao, Zhiyuan Yan, Haoming Xu, Yunzhen Zhao et al.CVPR 2026 · 7 citations
- Reasoning over Boundaries: Enhancing Specification Alignment via Test-time DeliberationHaoran Zhang, Yafu Li, Xuyang Hu, Dongrui Liu et al.ICML 2026 · 3 citations
- FLARE-AI: Flaw Reporting for AIShayne Longpre, Elaine Zhu, Carson Ezell, Avijit Ghosh et al.ICML 2026
- VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language ModelsWoojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung DoICML 2026
Builds on4
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang et al.ACL 2024 · 64 citations
Related papers
- SoSBench: Benchmarking Safety Alignment on Six Scientific DomainsFengqing Jiang, Fengbo Ma, Zhangchen Xu, Yuetai Li et al.ICLR 2026 · 14 citations
- YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language ModelsYaman Yu, Yiren Liu, Yuqi Zhang, Yun Huang et al.CCS 2025 · 1 citation
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalTinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang et al.ICLR 2025
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative ComprehensionQian Yang, Jin Xu, Wenrui Liu, Yunfei Chu et al.ACL 2024
