Lune

EMNLP2025顶会

Nullspace Disentanglement for Red Teaming Language Models

Yi Han, Yuanxing Liu, Weinan Zhang, Ting Liu

2025年份

摘要

With the widespread deployment of generative language models, concerns about safety issues have continuously grown. High-quality finetuning data generated from red teaming plays a crucial role in the model's safety. Recently, automated red teaming approaches have been proposed to create test cases. However, these approaches, which rely on open-ended generation, encounter issues related to inefficiency and low attack success rates. In this work, we introduce a black-box approach that ingeniously exploits the unique properties of the nullspace to disentangle and regulate the crucial success information within test cases. Our study provides a brand-new perspective for automated red team research. Experimental results demonstrate that our approach outperforms baseline methods regarding the attack success rate. The generated test cases also excel in aspects of diversity and fluency. Our code is available at: https://github.com/HITSCIR-DT-Code/NDR . Warning: Some examples shown in this paper can be offensive and upsetting. Rishabh Bhardwaj and Soujanya Poria. 2023. Redteaming large language models using chain of utterances for safety-alignment. ArXiv, abs/2308.09662.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖