PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
Yingjie Ma, Xun Lin, Yong Xu, Weicheng Xie, Zitong Yu
Abstract
In recent years, face anti-spoofing (FAS) has made notable progress in multimodal fusion, cross-domain generalization, and interpretability. With the development of large language models and reinforcement learning (RL), strategy-based training paradigms offer new opportunities for jointly modeling multimodality, generalization, and interpretability. However, compared to unimodal reasoning, multimodal reasoning introduces more complex logic, such as accurate feature representation and cross-modal verification, which significantly increases reasoning complexity and labeling difficulty. Due to the lack of high-quality annotations in existing multimodal FAS datasets, directly applying RL strategies is sub-optimal, hindering robust multimodal reasoning. In this paper, we find two key issues of supervised fine-tuning combined with reinforcement learning (SFT+RL) paradigms in multimodal FAS reasoning: 1) limited multimodal reasoning paths not only hinder the full utilization of multimodal information but also constrain the model’s exploration space after SFT, thereby affecting the effectiveness of subsequent RL; and 2) the mismatch between single-task supervision and the diversity of multimodal reasoning paths leads to reasoning confusion, where models may exploit shortcuts by directly mapping input images to answers, bypassing the intended reasoning process. These issues further increase the complexity of multimodal reasoning and hinder the effective application of RL strategies. To address these challenges, we propose the PA-FAS framework with a reasoning path enhancement strategy for high-quality extended reasoning sequences construction based on limited annotated data to enrich the reasoning paths and alleviate exploration constraints. Additionally, we introduce an answer shuffling mechanism during SFT for comprehensive multimodal analysis rather than mining superficial cues, thus encouraging deeper reasoning and avoiding shortcut learning. Our method significantly improves multimodal reasoning accuracy and generalization, and successfully unifies multimodal fusion, cross-domain generalization, and interpretability towards trustworthy multimodal FAS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d24b4b3-a73d-471a-aa43-4f88e183c841Builds on14
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- FLIP: Cross-domain Face Anti-spoofing with Language GuidanceKoushik Srivatsan, Muzammal Naseer, Karthik NandakumarICCV 2023 · 84 citations
- CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingAjian Liu, Shuai Xue, Jianwen Gan, Jun Wan et al.CVPR 2024 · 59 citations
- Detection and Continual Learning of Novel Face Presentation AttacksMohammad Rostami, Leonidas Spinoulas, Mohamed E. Hussein, Joe Mathai et al.ICCV 2021 · 51 citations
- Towards Unsupervised Domain Generalization for Face Anti-SpoofingYuchen Liu, Yabo Chen, Mengran Gou, Chun-Ting Huang et al.ICCV 2023 · 41 citations
Related papers
- From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-SpoofingHaoyuan Zhang, Keyao Wang, Guosheng Zhang, Haixiao Yue et al.CVPR 2026 · 2 citations
- Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation DetectionYuchen Zhang, Yaxiong Wang, Kecheng Han, Yujiao Wu et al.ACL 2026 · 1 citation
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish ViewJianyu Qi, Ding Zou, Wenrui Yan, Rui Ma et al.AAAI 2026
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
- Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-SpoofingHonglu Zhang, Zhiqin Fang, Ningning Zhao, Saihui Hou et al.CVPR 2026 · 4 citations
