LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
Hayder Elesedy, Pedro M. Esperança, Silviu Vlad Oprea, Mete Ozay
摘要
Guardrails have emerged as comprehensive method of content moderation for large language models (LLMs), complementing safety alignment from fine-tuning. However, existing model-based guardrails are too memory intensive for use on resource-constrained computational devices such as mobile phones, an increasing number of which are running LLM-based applications locally. We introduce LoRA-Guard, a parameter-efficient guardrail adaptation method that relies on knowledge sharing between LLMs and guardrail models. LoRA-Guard extracts language features from the LLMs and adapts them for the content moderation task using low-rank adapters in a dual-path design which prevents any performance degradation on the generative task. We show that LoRA-Guard outperforms existing guardrail approaches while using 100-1000x fewer guardrail parameters, enabling on-device content moderation. * Version Note: Changes in this version v2 relative to v1: separate output heads for safe/unsafe classification and harm category classification ( §4.3), training on BeaverTails dataset ( §4.2.1), use of recent chat models ( §4.1), comparison with recent guard models ( §5).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Shape it Up! Restoring LLM Safety during FinetuningShengyun Peng, Pin-Yu Chen, Jianfeng Chi, Seongmin Lee 等NeurIPS 2025 · 被引用 17 次
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim 等ICLR 2026 · 被引用 3 次
- A Game-Theoretic Analysis of Attacks on Large Language Models via Compositional SkillsXinbo Wu, Huan Zhang, Abhishek Umrawal, Lav VarshneyICML 2026
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
相关 Paper
- SaLoRA: Safety-Alignment Preserved Low-Rank AdaptationMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等ICLR 2025
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-SpaceBingjie Zhang, Yibo Yang, Renzhe, Dandan Guo 等ICLR 2026 · 被引用 12 次
- Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language ModelsLijia Lv, Yuanshu Zhao, Guan Wang, Xuehai Tang 等EMNLP 2025
- BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer SharingYuhua Zhou, Ruifeng Li, Changhai Zhou, Fei Yang 等ICML 2025
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth 等EMNLP 2025
