ORFuzz: Fuzzing the "Other Side" of LLM Safety - Testing Over-Refusal
Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Jiashui Wang, Xinlei Ying, Long Liu, Wenhai Wang
Abstract
Large Language Models (LLMs) have been found to show over-refusal problems—erroneously rejecting benign queries due to overly conservative safety measures—a critical functional flaw that undermines their reliability and usability. Current methods for testing this behavior are demonstrably inadequate, suffering from flawed benchmarks and limited test generation capabilities, as highlighted by our empirical user study. To the best of our knowledge, this paper introduces the first evolutionary testing framework, ORFuzz, for the systematic detection and analysis of LLM over-refusals. ORFuzz uniquely integrates three core components: (1) safety category-aware seed selection for comprehensive test coverage, (2) adaptive mutator optimization using reasoning LLMs to generate effective test cases, and (3) OR-Judge, a human-aligned judge model validated to accurately reflect user perception of toxicity and refusal. Our extensive evaluations demonstrate that ORFuzz generates diverse, validated over-refusal instances at a rate (6.98% average) more than double that of leading baselines, effectively uncovering vulnerabilities. Furthermore, ORFuzz’s outputs form the basis of ORFuzzSet, a new benchmark of 1,786 highly transferable test cases that achieves a superior 57.37% average over-refusal rate across 14 diverse LLMs, significantly outperforming existing datasets. ORFuzz and ORFuzzSet provide a robust automated testing framework and a valuable community resource, paving the way for developing more reliable and trustworthy LLM-based software systems. The code of this paper is available at: https://github.com/HotBento/ORFuzz.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 989ea662-ce16-4b5b-a01c-10a6f1430755Cited by top-tier papers2
- LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector AlignmentHaonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen et al.ACL 2026 · 1 citation
- DDOR: Delta Debugging for Explainable Overrefusal Testing and RepairQinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang et al.ISSTA 2026
Builds on4
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
- DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert KnowledgeBufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu et al.UbiComp 2025 · 60 citations
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language ModelsXiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen et al.ISSTA 2025 · 4 citations
- OR-Bench: An Over-Refusal Benchmark for Large Language ModelsJustin Cui, Wei-Lin Chiang, Ion Stoica, Cho-Jui HsiehICML 2025
Related papers
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious InstructionsXiaorui Wu, Fei Li, Xiaofeng Mao, Xin Zhang et al.NeurIPS 2025 · 8 citations
- Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision BoundaryLicheng Pan, Yongqi Tong, Xin Zhang, Xiaolu Zhang et al.EMNLP 2025 · 3 citations
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
- CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code ReasoningMan Ho Lam, Chaozheng Wang, Jen-Tse Huang, Michael R. LyuNeurIPS 2025 · 16 citations
- Red Teaming LLMs via Linguistic-Aware FuzzingShuai Yuan, Nian Luo, Jingling Sun, Yihao Huang et al.FSE 2026
