A Lightweight Explainable Guardrail for Prompt Safety
Md. Asiful Islam, Mihai Surdeanu
摘要
We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly learn a prompt classifier and an explanation classifier, where the latter labels prompt words that explain the safe/unsafe overall decision. LEG is trained on synthetic explanation data, which is generated using a novel strategy that counteracts the confirmation biases of LLMs. Lastly, LEG's training process uses a novel loss that captures global explanation signals as a weak supervision and combines crossentropy and focal losses with uncertainty-based weighting. LEG obtains equivalent or better performance than the state-of-the-art for both prompt classification and explainability, both in-domain and out-of-domain on three datasets, despite the fact that its model size is considerably smaller than current approaches. Code 1 Models and Datasets 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- A Causal Explainable Guardrails for Large Language ModelsZhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang 等CCS 2024 · 被引用 5 次
- Bootstrapping Language Models with DPO Implicit RewardsChangyu Chen, Zichen Liu, Chao Du, Tianyu Pang 等ICLR 2025
- InfAlign: Inference-aware language model alignmentAnanth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein 等ICML 2025
- R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical ReasoningMintong Kang, Bo LiICLR 2025
相关 Paper
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth 等EMNLP 2025
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim 等ICLR 2026 · 被引用 3 次
- Harnessing Hyperbolic Geometry for Harmful Prompt Detection and SanitizationIgor Maljkovic, Maria Rosaria Briglia, Iacopo Masi, Antonio Emanuele Cinà 等ICLR 2026 · 被引用 2 次
- BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric DebateArnon Mazza, Elad LeviICML 2026
- On Prompt-Driven Safeguarding for Large Language ModelsChujie Zheng, Fan Yin, Hao Zhou, Fandong Meng 等ICML 2024 · 被引用 116 次
