ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
Wonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu, Ashkan Yousefpour, Haon Park, Bumsub Ham, Suhyun Kim
Abstract
Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation methods, but these methods struggle to detect implicit harmful content or produce inaccurate evaluations. Therefore, we found that existing benchmarks have low levels of harmfulness, ambiguous data, and limited diversity in image-text pair combinations. To address these issues, we propose the ELITE benchmark, a high-quality safety evaluation benchmark for VLMs, underpinned by our enhanced evaluation method, the ELITE evaluator. The ELITE evaluator explicitly incorporates a toxicity score to accurately assess harmfulness in multimodal contexts, where VLMs often provide specific, convincing, but unharmful descriptions of images. We filter out ambiguous and low-quality image-text pairs from existing benchmarks using the ELITE evaluator and generate diverse combinations of safe and unsafe image-text pairs. Our experiments demonstrate that the ELITE evaluator achieves superior alignment with human evaluations compared to prior automated methods, and the ELITE benchmark offers enhanced benchmark quality and diversity. By introducing ELITE, we pave the way for safer, more robust VLMs, contributing essential tools for evaluating and mitigating safety risks in realworld applications. Warning: This paper includes examples of harmful language and images that may be sensitive or uncomfortable. Reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 060bcdd0-ecea-4877-86ca-0340d3e6c887Cited by top-tier papers4
- Jailbreaking on Text-to-Video Models via Scene Splitting StrategyWonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham et al.ICLR 2026 · 10 citations
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI SafetyShruti Palaskar, Leon Alexander Gatys, Mona Abdelrahman, Mar Jacobo et al.ICLR 2026 · 8 citations
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMsDasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt et al.ACL 2026 · 3 citations
- The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image ReasoningRenmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang et al.ACL 2026 · 1 citation
Builds on3
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang et al.ICML 2024 · 140 citations
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
Related papers
- Iterative Prompt Refinement for Safer Text-to-Image GenerationJinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun LeeEMNLP 2025
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang et al.ACL 2025
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language ModelsHanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang et al.AAAI 2026 · 2 citations
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes et al.CCS 2025 · 1 citation
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong et al.ACL 2026 · 10 citations
