Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models
Lijia Lv, Yuanshu Zhao, Guan Wang, Xuehai Tang, Jie Wen, Jizhong Han, Songlin Hu
Abstract
Disclaimer: This paper contains potentially offensive and harmful text. Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. Yet studies show that guard models (Guards)-LLMs fine-tuned for safety-remain vulnerable to "jailbreak" attacks, jeopardising downstream chatbots. We confirm this weakness on three public benchmarks (BeaverTails, XSTest, AdvBench) and trace it to representation shifts that arise in the embedding layer and cascade through the Transformer stack. To counteract the effect, we introduce Gamma-Guard: lightweight residual adapters inserted after the embeddings and at sparse intervals in the model. The adapters start with zero-scaled gates, so they retain the original behaviour; a brief adversarial finetuning phase then teaches them to denoise embeddings and refocus attention. With fewer than 0.1 % extra parameters and only a 2 % latency increase, Gamma-Guard lifts adversarial accuracy from ≤ 5 % to ≈ 95 % a 90 percentage-point gain while reducing cleandata accuracy by just 8 percentage points. Extensive ablations further show that robustness improvements persist across different layer placements and model sizes. To our knowledge, this is the first approach that directly augments large Guards with trainable adapters, providing a practical path toward safer large-scale LLM deployments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54737b5b-cd50-47d9-b255-8ea72e39dc83Cited by top-tier papers1
Ask how each one uses itBuilds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
Related papers
- LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language ModelsHayder Elesedy, Pedro M. Esperança, Silviu Vlad Oprea, Mete OzayEMNLP 2024 · 5 citations
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi Xie, Minghong Fang, Renjie Pi, Neil GongACL 2024
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song et al.ACL 2026 · 22 citations
- SAFENUDGE: Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offsJoão Fonseca, Andrew Bell, Julia StoyanovichEMNLP 2025 · 1 citation
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang et al.ICML 2024 · 140 citations
