Lune

EMNLP2024Top-tier venue

LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models

Hayder Elesedy, Pedro M. Esperança, Silviu Vlad Oprea, Mete Ozay

2024Year
5Citations
3Top-tier citations

Abstract

Guardrails have emerged as comprehensive method of content moderation for large language models (LLMs), complementing safety alignment from fine-tuning. However, existing model-based guardrails are too memory intensive for use on resource-constrained computational devices such as mobile phones, an increasing number of which are running LLM-based applications locally. We introduce LoRA-Guard, a parameter-efficient guardrail adaptation method that relies on knowledge sharing between LLMs and guardrail models. LoRA-Guard extracts language features from the LLMs and adapts them for the content moderation task using low-rank adapters in a dual-path design which prevents any performance degradation on the generative task. We show that LoRA-Guard outperforms existing guardrail approaches while using 100-1000x fewer guardrail parameters, enabling on-device content moderation. * Version Note: Changes in this version v2 relative to v1: separate output heads for safe/unsafe classification and harm category classification ( §4.3), training on BeaverTails dataset ( §4.2.1), use of recent chat models ( §4.1), comparison with recent guard models ( §5).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 50f6836e-ccfd-439e-84ab-d091016aedbc

Cited by top-tier papers3

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines