Lune

EMNLP2025Top-tier venue

Nullspace Disentanglement for Red Teaming Language Models

Yi Han, Yuanxing Liu, Weinan Zhang, Ting Liu

2025Year

Abstract

With the widespread deployment of generative language models, concerns about safety issues have continuously grown. High-quality finetuning data generated from red teaming plays a crucial role in the model's safety. Recently, automated red teaming approaches have been proposed to create test cases. However, these approaches, which rely on open-ended generation, encounter issues related to inefficiency and low attack success rates. In this work, we introduce a black-box approach that ingeniously exploits the unique properties of the nullspace to disentangle and regulate the crucial success information within test cases. Our study provides a brand-new perspective for automated red team research. Experimental results demonstrate that our approach outperforms baseline methods regarding the attack success rate. The generated test cases also excel in aspects of diversity and fluency. Our code is available at: https://github.com/HITSCIR-DT-Code/NDR . Warning: Some examples shown in this paper can be offensive and upsetting. Rishabh Bhardwaj and Soujanya Poria. 2023. Redteaming large language models using chain of utterances for safety-alignment. ArXiv, abs/2308.09662.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5b0406d0-97f7-42f8-b2de-38aeeb71eb07

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines