Anecdoctoring: Automated Red-Teaming Across Language and Place
Alejandro Cuevas, Saloni Dash, Bharat Kumar Nayak, Dan Vann, Madeleine I. G. Daepp
摘要
Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates redteaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages and cultures, but red-teaming datasets are commonly US-and English-centric. To address this gap, we propose "anecdoctoring", a novel red-teaming approach that automatically generates adversarial prompts across languages and cultures. We collect misinformation claims from fact-checking websites in three languages (English, Spanish, and Hindi) and two geographies (US and India). We then cluster individual claims into broader narratives and characterize the resulting clusters with knowledge graphs, with which we augment an attacker LLM. Our method produces higher attack success rates and offers interpretability benefits relative to few-shot prompting. Results underscore the need for disinformation mitigations that scale globally and are grounded in realworld adversarial misuse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 被引用 230 次
相关 Paper
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce HarmAakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant 等EMNLP 2024 · 被引用 3 次
- To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMsZohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi 等ACL 2026
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM AgentsNivya Talokar, Ayush Kumar Tarun, Murari Mandal, Maksym Andriushchenko 等ICML 2026 · 被引用 2 次
- CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark GenerationChaeyun Kim, YongTaek Lim, Kihyun Kim, Junghwan Kim 等ICLR 2026 · 被引用 2 次
