FLIP: Cross-domain Face Anti-spoofing with Language Guidance
Koushik Srivatsan, Muzammal Naseer, Karthik Nandakumar
Abstract
Face anti-spoofing (FAS) or presentation attack detection is an essential component of face recognition systems deployed in security-critical applications. Existing FAS methods have poor generalizability to unseen spoof types, camera sensors, and environmental conditions. Recently, vision transformer (ViT) models have been shown to be effective for the FAS task due to their ability to capture long-range dependencies among image patches. However, adaptive modules or auxiliary loss functions are often required to adapt pre-trained ViT weights learned on large-scale datasets such as ImageNet. In this work, we first show that initializing ViTs with multimodal (e.g., CLIP) pre-trained weights improves generalizability for the FAS task, which is in line with the zero-shot transfer capabilities of vision-language pre-trained (VLP) models. We then propose a novel approach for robust cross-domain FAS by grounding visual representations with the help of natural language. Specifically, we show that aligning the image representation with an ensemble of class descriptions (based on natural language semantics) improves FAS generalizability in low-data regimes. Finally, we propose a multimodal contrastive learning strategy to boost feature generalization further and bridge the gap between source and target domains. Extensive experiments on three standard protocols demonstrate that our method significantly outperforms the state-of-the-art methods, achieving better zero-shot transfer performance than five-shot transfer of "adaptive ViTs". Code: https://github.com/koushiksrivats/FLIP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingAjian Liu, Shuai Xue, Jianwen Gan, Jun Wan et al.CVPR 2024 · 59 citations
- FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingAjian Liu, Hui Ma, Junze Zheng, Haocheng Yuan et al.ACM MM 2024 · 34 citations
- SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image PretrainingPei-Kai Huang, Jun-Xiong Chong, Cheng-Hsuan Chiang, Tzu-Hsien Chen et al.AAAI 2025 · 16 citations
- Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language ModelsGuosheng Zhang, Keyao Wang, Haixiao Yue, Ajian Liu et al.AAAI 2025 · 13 citations
- mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionGeng Chen, Wuyuan Xie, Di Lin, Ye Liu et al.AAAI 2025 · 7 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Multi-View Slot Attention using Paraphrased Texts for Face Anti-SpoofingJeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon et al.ICCV 2025 · 5 citations
- Style-conditional Prompt Token Learning for Generalizable Face Anti-spoofingJiabao Guo, Huan Liu, Yizhi Luo, Xueli Hu et al.ACM MM 2024 · 18 citations
- Fine-Grained Prompt Learning for Face Anti-SpoofingXueli Hu, Huan Liu, Haocheng Yuan, Zhiyang Fu et al.ACM MM 2024 · 9 citations
- InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofingKun-Hsiang Lin, Yu-Wen Tseng, Kang-Yang Huang, Jhih-Ciang Wu et al.ACM MM 2025 · 4 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
