A Lightweight Explainable Guardrail for Prompt Safety
Md. Asiful Islam, Mihai Surdeanu
Abstract
We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly learn a prompt classifier and an explanation classifier, where the latter labels prompt words that explain the safe/unsafe overall decision. LEG is trained on synthetic explanation data, which is generated using a novel strategy that counteracts the confirmation biases of LLMs. Lastly, LEG's training process uses a novel loss that captures global explanation signals as a weak supervision and combines crossentropy and focal losses with uncertainty-based weighting. LEG obtains equivalent or better performance than the state-of-the-art for both prompt classification and explainability, both in-domain and out-of-domain on three datasets, despite the fact that its model size is considerably smaller than current approaches. Code 1 Models and Datasets 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 586dcc1f-b5fe-4197-8fc2-7a2efeffab98Builds on6
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- A Causal Explainable Guardrails for Large Language ModelsZhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang et al.CCS 2024 · 5 citations
- Bootstrapping Language Models with DPO Implicit RewardsChangyu Chen, Zichen Liu, Chao Du, Tianyu Pang et al.ICLR 2025
- InfAlign: Inference-aware language model alignmentAnanth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein et al.ICML 2025
- R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical ReasoningMintong Kang, Bo LiICLR 2025
Related papers
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth et al.EMNLP 2025
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim et al.ICLR 2026 · 3 citations
- Harnessing Hyperbolic Geometry for Harmful Prompt Detection and SanitizationIgor Maljkovic, Maria Rosaria Briglia, Iacopo Masi, Antonio Emanuele Cinà et al.ICLR 2026 · 2 citations
- BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric DebateArnon Mazza, Elad LeviICML 2026
- On Prompt-Driven Safeguarding for Large Language ModelsChujie Zheng, Fan Yin, Hao Zhou, Fandong Meng et al.ICML 2024 · 116 citations
