FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors
Sepehr Dehdashtian, Jacob Seidman, Vishnu Boddeti, Gaurav Bharaj
Abstract
Audio deepfake detection (ADD) models are critical for countering the malicious use of text-tospeech (TTS) models. Evaluating and strengthening ADD models requires developing datasets that span the space of generated audio and highlight high-error regions. Existing dataset development strategies face two challenges: (i) manual collection, and (ii) inefficient discovery of blind spots in the ADD models. To address these challenges, we propose FoeGlass, the first black-box automated red-teaming method for ADDs, which effectively discovers ADD failure modes in the space of generated audio underexplored by state-of-theart deepfake benchmarks. FoeGlass uses the incontext learning capabilities of an LLM to explore the input space of a TTS model, generating audio samples that fool the target ADD using only black-box access to all components. By using a carefully designed context based on diversity measurements, FoeGlass mitigates the common problem of mode collapse in automated red-teaming systems. Empirical evaluations on several opensource ADD and TTS models demonstrate that data generated from FoeGlass substantially improves the false negative rates over unconditional sampling baselines and recent spoofing datasets by up to 94%, while requiring no manual supervision. Furthermore, we show that the attacks generated by FoeGlass are transferable across different target ADDs, demonstrating its broad applicability and ease of use for the automated red teaming of ADD systems. Finally, fine-tuning ADD models on FoeGlass-generated samples notably enhances the robustness of the detectors (up to 41%). †This work was done during an internship at Reality Defender.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
Related papers
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image DetectorsSepehr Dehdashtian, Mashrur Mahmud Morshed, Jacob H. Seidman, Gaurav Bharaj et al.NeurIPS 2025 · 1 citation
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 29 citations
- VoiceWukong: Benchmarking Deepfake Voice DetectionZiwei Yan, Yanjie Zhao, Haoyu WangUSENIX Security 2025
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang et al.WWW 2023 · 34 citations
- Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake DetectionYuze Zhao, Kuiyuan Zhang, Zhongyun Hua, Yushu Zhang et al.ICML 2026
