Evading Forensic Classifiers with Attribute-Conditioned Adversarial Faces
Fahad Shamshad, Koushik Srivatsan, Karthik Nandakumar
Abstract
The ability of generative models to produce highly realistic synthetic face images has raised security and ethical concerns. As a first line of defense against such fake faces, deep learning based forensic classifiers have been developed. While these forensic models can detect whether a face image is synthetic or real with high accuracy, they are also vulnerable to adversarial attacks. Although such attacks can be highly successful in evading detection by forensic classifiers, they introduce visible noise patterns that are detectable through careful human scrutiny. Additionally, these attacks assume access to the target model(s) which may not always be true. Attempts have been made to directly perturb the latent space of GANs to produce adversarial fake faces that can circumvent forensic classifiers. In this work, we go one step further and show that it is possible to successfully generate adversarial fake faces with a specified set of attributes (e.g., hair color, eye size, race, gender, etc.). To achieve this goal, we leverage the state-of-the-art generative model StyleGAN with disentangled representations, which enables a range of modifications without leaving the manifold of natural images. We propose a framework to search for adversarial latent codes within the feature space of StyleGAN, where the search can be guided either by a text prompt or a reference image. We also propose a metalearning based optimization strategy to achieve transferable performance on unknown target models. Extensive experiments demonstrate that the proposed approach can produce semantically manipulated adversarial fake faces, which are true to the specified attribute set and can successfully fool forensic face classifiers, while remaining undetectable by humans. Code: https://github.com/ koushiksrivats/face_attribute_attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dbaa00d-8953-4fb5-b0c2-83c57df1ac32Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Attributing Fake Images to GANs: Learning and Analyzing GAN FingerprintsNing Yu, Larry Davis, Mario FritzICCV 2019 · 533 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
Related papers
- Exploring Adversarial Fake Images on Face ManifoldDongze Li, Wei Wang, Hongxing Fan, Jing DongCVPR 2021
- ImU: Physical Impersonating Attack for Face Recognition System with Natural Style ChangesShengwei An, Yuan Yao, Qiuling Xu, Shiqing Ma et al.S&P 2023
- Everything is There in Latent Space: Attribute Editing and Attribute Style Manipulation by StyleGAN Latent Space ExplorationRishubh Parihar, Ankit Dhiman, Tejan Karmali, Venkatesh Babu R.ACM MM 2022 · 21 citations
- Attribute-specific Control Units in StyleGAN for Fine-grained Image ManipulationRui Wang, Jian Chen, Gang Yu, Li Sun et al.ACM MM 2021 · 13 citations
- Adaptive Nonlinear Latent Transformation for Conditional Face EditingZhizhong Huang, Siteng Ma, Junping Zhang, Hongming ShanICCV 2023 · 13 citations
