mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention Fusion
Geng Chen, Wuyuan Xie, Di Lin, Ye Liu, Miaohui Wang
Abstract
The increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e6010c9-507a-48d1-95b9-8f33268316dfBuilds on10
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- FLIP: Cross-domain Face Anti-spoofing with Language GuidanceKoushik Srivatsan, Muzammal Naseer, Karthik NandakumarICCV 2023 · 84 citations
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 74 citations
- Adaptive Mixture of Experts Learning for Generalizable Face Anti-SpoofingQianyu Zhou, Ke-Yue Zhang, Taiping Yao, Ran Yi et al.ACM MM 2022 · 65 citations
- FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingAjian Liu, Hui Ma, Junze Zheng, Haocheng Yuan et al.ACM MM 2024 · 34 citations
Related papers
- DADM: Dual Alignment of Domain and Modality for Face Anti-SpoofingJingyi Yang, Xun Lin, Zitong Yu, Liepiao Zhang et al.ICCV 2025 · 1 citation
- Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-SpoofingXun Lin, Shuai Wang, Rizhao Cai, Yizhong Liu et al.CVPR 2024
- Deep Spatial Gradient and Temporal Depth Learning for Face Anti-SpoofingZezheng Wang, Zitong Yu, Chenxu Zhao, Xiangyu Zhu et al.CVPR 2020
- Cross Modal Focal Loss for RGBD Face Anti-SpoofingAnjith George, Sébastien MarcelCVPR 2021
- Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReIDDe Cheng, Lingfeng He, Nannan Wang, Shizhou Zhang et al.ACM MM 2023 · 36 citations
