Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
Guosheng Zhang, Keyao Wang, Haixiao Yue, Ajian Liu, Gang Zhang, Kun Yao, Errui Ding, Jingdong Wang
摘要
Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They exhibit limited generalization in out-of-domain scenarios, such as new environments or unseen spoofing types. In this work, we introduce a multimodal large language model (MLLM) framework for FAS, termed Interpretable Face Anti-Spoofing (I-FAS), which transforms the FAS task into an interpretable visual question answering (VQA) paradigm. Specifically, we propose a Spoof-aware Captioning and Filtering (SCF) strategy to generate high-quality captions for FAS images, enriching the model's supervision with natural language interpretations. To mitigate the impact of noisy captions during training, we develop a Lopsided Language Model (L-LM) loss function that separates loss calculations for judgment and interpretation, prioritizing the optimization of the former. Furthermore, to enhance the model's perception of global visual features, we design a Globally Aware Connector (GAC) to align multi-level visual representations with the language model. Extensive experiments on standard and newly devised One to Eleven cross-domain benchmarks, comprising 12 public datasets, demonstrate that our method significantly outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-SpoofingHonglu Zhang, Zhiqin Fang, Ningning Zhao, Saihui Hou 等CVPR 2026 · 被引用 4 次
- InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofingKun-Hsiang Lin, Yu-Wen Tseng, Kang-Yang Huang, Jhih-Ciang Wu 等ACM MM 2025 · 被引用 4 次
- From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-SpoofingHaoyuan Zhang, Keyao Wang, Guosheng Zhang, Haixiao Yue 等CVPR 2026 · 被引用 2 次
- PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement LearningYingjie Ma, Xun Lin, Yong Xu, Weicheng Xie 等AAAI 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Domain Generalization via Shuffled Style Assembly for Face Anti-SpoofingZhuo Wang, Zezheng Wang, Zitong Yu, Weihong Deng 等CVPR 2022 · 被引用 195 次
相关 Paper
- MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLMTao Chen, Jingyi Zhang, Decheng Liu, Chunlei PengWWW 2026 · 被引用 1 次
- FLIP: Cross-domain Face Anti-spoofing with Language GuidanceKoushik Srivatsan, Muzammal Naseer, Karthik NandakumarICCV 2023 · 被引用 84 次
- Multi-View Slot Attention using Paraphrased Texts for Face Anti-SpoofingJeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon 等ICCV 2025 · 被引用 5 次
- FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingAjian Liu, Hui Ma, Junze Zheng, Haocheng Yuan 等ACM MM 2024 · 被引用 34 次
- UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationQihui Zhang, Munan Ning, Zheyuan Liu, Yue Huang 等CVPR 2025
