OSCS: Online Selection with Provable FAR Control for LLM Safety
Zirui Hu, Zheng Zhang, Yingjie Wang, Dacheng Tao
Abstract
Large language models (LLMs) are vulnerable to malicious inputs, posing serious risks in high-stakes applications. Although existing detection-based defenses achieve strong empirical performance, they generally lack explicit control over the false acceptance rate (FAR), a critical safety requirement in sensitive deployment scenarios. This challenge is further complicated by two practical constraints: the lack of malicious calibration samples and the streaming nature of real-world inputs. To address these challenges, we propose OSCS, a novel framework for online FAR control without requiring malicious calibration data. OSCS leverages detection scores produced by existing defenses and employs recursive density estimation to estimate benign probability from the test stream. Based on these estimates, OSCS performs real-time accept/reject decisions while provably satisfying a user-specified FAR target. Theoretically, we show that OSCS controls the FAR up to a vanishing excess term under mild assumptions. Extensive experiments on backdoor and jailbreak attack tasks further demonstrate the effectiveness of OSCS, showing that it consistently achieves robust FAR control across diverse attack settings while outperforming existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- On Prompt-Driven Safeguarding for Large Language ModelsChujie Zheng, Fan Yin, Hao Zhou, Fandong Meng et al.ICML 2024 · 116 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP ModelsWenkai Yang, Yankai Lin, Peng Li, Jie Zhou et al.EMNLP 2021 · 57 citations
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu et al.NeurIPS 2023 · 18 citations
Related papers
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou et al.ACL 2026 · 4 citations
- ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsZihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan et al.AAAI 2026 · 5 citations
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming GuardrailsMintong Kang, Zhaorun Chen, Bo LiNeurIPS 2025 · 4 citations
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language ModelsGuangyu Yang, Jinghong Chen, Jingbiao Mei, Weizhe Lin et al.ACL 2026 · 1 citation
- LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed EffectYixiao Xu, Binxing Fang, Mohan Li, Keke Tang et al.NeurIPS 2024 · 7 citations
