FLIP: Cross-domain Face Anti-spoofing with Language Guidance
Koushik Srivatsan, Muzammal Naseer, Karthik Nandakumar
摘要
Face anti-spoofing (FAS) or presentation attack detection is an essential component of face recognition systems deployed in security-critical applications. Existing FAS methods have poor generalizability to unseen spoof types, camera sensors, and environmental conditions. Recently, vision transformer (ViT) models have been shown to be effective for the FAS task due to their ability to capture long-range dependencies among image patches. However, adaptive modules or auxiliary loss functions are often required to adapt pre-trained ViT weights learned on large-scale datasets such as ImageNet. In this work, we first show that initializing ViTs with multimodal (e.g., CLIP) pre-trained weights improves generalizability for the FAS task, which is in line with the zero-shot transfer capabilities of vision-language pre-trained (VLP) models. We then propose a novel approach for robust cross-domain FAS by grounding visual representations with the help of natural language. Specifically, we show that aligning the image representation with an ensemble of class descriptions (based on natural language semantics) improves FAS generalizability in low-data regimes. Finally, we propose a multimodal contrastive learning strategy to boost feature generalization further and bridge the gap between source and target domains. Extensive experiments on three standard protocols demonstrate that our method significantly outperforms the state-of-the-art methods, achieving better zero-shot transfer performance than five-shot transfer of "adaptive ViTs". Code: https://github.com/koushiksrivats/FLIP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingAjian Liu, Shuai Xue, Jianwen Gan, Jun Wan 等CVPR 2024 · 被引用 59 次
- FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingAjian Liu, Hui Ma, Junze Zheng, Haocheng Yuan 等ACM MM 2024 · 被引用 34 次
- SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image PretrainingPei-Kai Huang, Jun-Xiong Chong, Cheng-Hsuan Chiang, Tzu-Hsien Chen 等AAAI 2025 · 被引用 16 次
- Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language ModelsGuosheng Zhang, Keyao Wang, Haixiao Yue, Ajian Liu 等AAAI 2025 · 被引用 13 次
- mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionGeng Chen, Wuyuan Xie, Di Lin, Ye Liu 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Multi-View Slot Attention using Paraphrased Texts for Face Anti-SpoofingJeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon 等ICCV 2025 · 被引用 5 次
- Style-conditional Prompt Token Learning for Generalizable Face Anti-spoofingJiabao Guo, Huan Liu, Yizhi Luo, Xueli Hu 等ACM MM 2024 · 被引用 18 次
- Fine-Grained Prompt Learning for Face Anti-SpoofingXueli Hu, Huan Liu, Haocheng Yuan, Zhiyang Fu 等ACM MM 2024 · 被引用 9 次
- InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofingKun-Hsiang Lin, Yu-Wen Tseng, Kang-Yang Huang, Jhih-Ciang Wu 等ACM MM 2025 · 被引用 4 次
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
