S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Jialuo Chen, Hui Xue, Xiaoxia Liu, Wenhai Wang, Kui Ren, Jingyi Wang
Abstract
Generative large language models (LLMs) have revolutionized natural language processing with their transformative and emergent capabilities. However, recent evidence indicates that LLMs can produce harmful content that violates social norms, raising significant concerns regarding the safety and ethical ramifications of deploying these advanced models. Thus, it is both critical and imperative to perform a rigorous and comprehensive safety evaluation of LLMs before deployment. Despite this need, owing to the extensiveness of LLM generation space, it still lacks a unified and standardized risk taxonomy to systematically reflect the LLM content safety, as well as automated safety assessment techniques to explore the potential risks efficiently. To bridge the striking gap, we propose S-Eval, a novel LLM-based automated Safety Evaluation framework with a newly defined comprehensive risk taxonomy. S-Eval incorporates two key components, i.e., an expert testing LLM M t and a novel safety critique LLM M c . The expert testing LLM M t is responsible for automatically generating test cases in accordance with the proposed risk management (including 8 risk dimensions and a total of 102 subdivided risks). The safety critique LLM M c can provide quantitative and explainable safety evaluations for better risk awareness of LLMs. In contrast to prior works, S-Eval differs in significant ways: (i) efficient – we construct a multi-dimensional and open-ended benchmark comprising 220,000 test cases across 102 risks utilizing M t and conduct safety evaluations for 21 influential LLMs via M c on our benchmark. The entire process is fully automated and requires no human involvement. (ii) effective – extensive validations show S-Eval facilitates a more thorough assessment and better perception of potential LLM risks, and M c not only accurately quantifies the risks of LLMs but also provides explainable and in-depth insights into their safety, surpassing comparable models such as LLaMA-Guard-2. (iii) adaptive – S-Eval can be flexibly configured and adapted to the rapid evolution of LLMs and accompanying new safety threats, test generation methods and safety critique methods thanks to the LLM-based architecture. We further study the impact of hyper-parameters and language environments on model safety, which may lead to promising directions for future research. S-Eval has been deployed in our industrial partner for the automated safety evaluation of multiple LLMs serving millions of users, demonstrating its effectiveness in real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f18f40e-de67-4b85-8a00-8f459f2749fdCited by top-tier papers10
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent ApproachYuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian et al.NeurIPS 2025 · 20 citations
- PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk DetectionXiaodan Li, Mengjie Wu, Yao Zhu, Yunna Lv et al.ICML 2026 · 5 citations
- LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector AlignmentHaonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen et al.ACL 2026 · 1 citation
- Adaptively profiling models with task elicitationDavis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar et al.EMNLP 2025 · 1 citation
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu et al.ICML 2026
Builds on8
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu et al.ACL 2024 · 9 citations
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
Related papers
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language ModelsHanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang et al.AAAI 2026 · 2 citations
- HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM SimulationYiran Liu, Zhiyi Hou, Xiaoang Xu, Shuo Wang et al.AAAI 2026
- LongSafety: Evaluating Long-Context Safety of Large Language ModelsYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui et al.ACL 2025 · 6 citations
- Dynamic Evaluation with Cognitive Reasoning for Multi-turn Safety of Large Language ModelsLanxue Zhang, Yanan Cao, Yuqiang Xie, Fang Fang et al.ACL 2025
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng et al.ICLR 2026 · 28 citations
