Lune

EMNLP2025顶会

Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models

Lijia Lv, Yuanshu Zhao, Guan Wang, Xuehai Tang, Jie Wen, Jizhong Han, Songlin Hu

2025年份
1顶会引用

摘要

Disclaimer: This paper contains potentially offensive and harmful text. Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. Yet studies show that guard models (Guards)-LLMs fine-tuned for safety-remain vulnerable to "jailbreak" attacks, jeopardising downstream chatbots. We confirm this weakness on three public benchmarks (BeaverTails, XSTest, AdvBench) and trace it to representation shifts that arise in the embedding layer and cascade through the Transformer stack. To counteract the effect, we introduce Gamma-Guard: lightweight residual adapters inserted after the embeddings and at sparse intervals in the model. The adapters start with zero-scaled gates, so they retain the original behaviour; a brief adversarial finetuning phase then teaches them to denoise embeddings and refocus attention. With fewer than 0.1 % extra parameters and only a 2 % latency increase, Gamma-Guard lifts adversarial accuracy from ≤ 5 % to ≈ 95 % a 90 percentage-point gain while reducing cleandata accuracy by just 8 percentage points. Extensive ablations further show that robustness improvements persist across different layer placements and model sizes. To our knowledge, this is the first approach that directly augments large Guards with trainable adapters, providing a practical path toward safer large-scale LLM deployments.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖