Nullspace Disentanglement for Red Teaming Language Models
Yi Han, Yuanxing Liu, Weinan Zhang, Ting Liu
Abstract
With the widespread deployment of generative language models, concerns about safety issues have continuously grown. High-quality finetuning data generated from red teaming plays a crucial role in the model's safety. Recently, automated red teaming approaches have been proposed to create test cases. However, these approaches, which rely on open-ended generation, encounter issues related to inefficiency and low attack success rates. In this work, we introduce a black-box approach that ingeniously exploits the unique properties of the nullspace to disentangle and regulate the crucial success information within test cases. Our study provides a brand-new perspective for automated red team research. Experimental results demonstrate that our approach outperforms baseline methods regarding the attack success rate. The generated test cases also excel in aspects of diversity and fluency. Our code is available at: https://github.com/HITSCIR-DT-Code/NDR . Warning: Some examples shown in this paper can be offensive and upsetting. Rishabh Bhardwaj and Soujanya Poria. 2023. Redteaming large language models using chain of utterances for safety-alignment. ArXiv, abs/2308.09662.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b0406d0-97f7-42f8-b2de-38aeeb71eb07Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving SupportMiriah Steiger, Timir J. Bharucha, Sukrit Venkatagiri, Martin J. Riedl et al.CHI 2021 · 168 citations
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton et al.ACL 2020 · 25 citations
Related papers
- Query-Efficient Black-Box Red Teaming via Bayesian OptimizationDeokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim et al.ACL 2023 · 5 citations
- Red Teaming LLMs via Linguistic-Aware FuzzingShuai Yuan, Nian Luo, Jingling Sun, Yihao Huang et al.FSE 2026
- Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language ModelsKai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang et al.ACL 2026
- Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn InteractionJinchuan Zhang, Yan Zhou, Yaxin Liu, Ziming Li et al.EMNLP 2024 · 2 citations
- Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and ActivationsRima Hazra, Sayan Layek, Somnath Banerjee, Soujanya PoriaEMNLP 2024
