Flow-Attention-based Spatio-Temporal Aggregation Network for 3D Mask Detection
Yuxin Cao, Yian Li, Yumeng Zhu, Derui Wang, Minhui Xue
Abstract
Anti-spoofing detection has become a necessity for face recognition systems due to the security threat posed by spoofing attacks. Despite great success in traditional attacks, most deep-learning-based methods perform poorly in 3D masks, which can highly simulate real faces in appearance and structure, suffering generalizability insufficiency while focusing only on the spatial domain with single frame input. This has been mitigated by the recent introduction of a biomedical technology called rPPG (remote photoplethysmography). However, rPPG-based methods are sensitive to noisy interference and require at least one second (> 25 frames) of observation time, which induces high computational overhead. To address these challenges, we propose a novel 3D mask detection framework, called FASTEN (Flow-Attention-based Spatio-Temporal aggrEgation Network). We tailor the network for focusing more on fine-grained details in large movements, which can eliminate redundant spatio-temporal feature interference and quickly capture splicing traces of 3D masks in fewer frames. Our proposed network contains three key modules: 1) a facial optical flow network to obtain non-RGB inter-frame flow information; 2) flow attention to assign different significance to each frame; 3) spatio-temporal aggregation to aggregate high-level spatial features and temporal transition features. Through extensive experiments, FASTEN only requires five frames of input and outperforms eight competitors for both intra-dataset and crossdataset evaluations in terms of multiple detection metrics. Moreover, FASTEN has been deployed in real-world mobile devices for practical 3D mask detection. * This work was done while Yuxin and Yian worked as interns at Ping An Technology (Shenzhen) Co., Ltd. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77263414-a12c-4695-9e36-45f6d8b3c12aBuilds on5
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video EnhancementZitong Yu, Wei Peng, Xiaobai Li, Xiaopeng Hong et al.ICCV 2019 · 324 citations
- Stealthy Adversarial Perturbations Against Real-Time Video Classification SystemsShasha Li, Ajaya Neupane, Sujoy Paul, Chengyu Song et al.NDSS 2019 · 132 citations
- StyleFool: Fooling Video Classification Systems via Style TransferYuxin Cao, Xi Xiao, Ruoxi Sun, Derui Wang et al.S&P 2023
- Searching Central Difference Convolutional Networks for Face Anti-SpoofingZitong Yu, Chenxu Zhao, Zezheng Wang, Yunxiao Qin et al.CVPR 2020
Related papers
- Deep Spatial Gradient and Temporal Depth Learning for Face Anti-SpoofingZezheng Wang, Zitong Yu, Chenxu Zhao, Xiangyu Zhu et al.CVPR 2020
- RhythmMamba: Fast, Lightweight, and Accurate Remote Physiological MeasurementBochao Zou, Zizheng Guo, Xiaocheng Hu, Huimin MaAAAI 2025 · 24 citations
- DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsHua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie et al.ACM MM 2020 · 224 citations
- Regularized Fine-Grained Meta Face Anti-SpoofingRui Shao, Xiangyuan Lan, Pong C. YuenAAAI 2020 · 185 citations
- Rhythmguassian: Repurposing Generalizable Gaussian Model for Remote Physiological MeasurementHao Lu, Yuting Zhang, Jiaqi Tang, Bowen Fu et al.ICCV 2025 · 2 citations
