Query-Efficient Black-Box Red Teaming via Bayesian Optimization
Deokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, Hyun Oh Song
摘要
The deployment of large-scale generative models is often restricted by their potential risk of causing harm to users in unpredictable ways. We focus on the problem of black-box red teaming, where a red team generates test cases and interacts with the victim model to discover a diverse set of failures with limited query access. Existing red teaming methods construct test cases based on human supervision or language model (LM) and query all test cases in a brute-force manner without incorporating any information from past evaluations, resulting in a prohibitively large number of queries. To this end, we propose Bayesian red teaming (BRT), novel query-efficient blackbox red teaming methods based on Bayesian optimization, which iteratively identify diverse positive test cases leading to model failures by utilizing the pre-defined user input pool and the past evaluations. Experimental results on various user input pools demonstrate that our method consistently finds a significantly larger number of diverse positive test cases under the limited query budget than the baseline methods. The source code is available at https://github.com/snu-mllab/Bayesian-Red-Teaming .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin 等ICLR 2024 · 被引用 207 次
- Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic PromptsZhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang, Pin-Yu Chen 等ICML 2024 · 被引用 155 次
- Curiosity-driven Red-teaming for Large Language ModelsZhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang 等ICLR 2024 · 被引用 84 次
- Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function PriorShuyu Cheng, Yibo Miao, Yinpeng Dong, Xiao Yang 等ICML 2024 · 被引用 15 次
- DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing ConstraintsAndrew Zhao, Quentin Xu, Matthieu Lin, Shenzhi Wang 等AAAI 2025 · 被引用 11 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
相关 Paper
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- BayesOpt Adversarial AttackBinxin Ru, Adam D. Cobb, Arno Blaas, Yarin GalICLR 2020 · 被引用 85 次
- Nullspace Disentanglement for Red Teaming Language ModelsYi Han, Yuanxing Liu, Weinan Zhang, Ting LiuEMNLP 2025
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA ExpertsChristos Ziakas, Nicholas Loo, Nishita Jain, Alessandra RussoACL 2026 · 被引用 2 次
- FLIRT: Feedback Loop In-context Red TeamingNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu 等EMNLP 2024 · 被引用 8 次
