Spatial Alignment and Temporal Matching Adapter for Video-Radar Remote Physiological Measurement
Qian Liang, Ruixu Geng, Jinbo Chen, Haoyu Wang, Yan Chen, Yang Hu
Abstract
Remote physiological measurement (RPM) based on video and radar has made significant progress in recent years. However, unimodal methods based solely on video or radar sensor have notable limitations due to their measurement principles, and multimodal RPM that combines these modalities has emerged as a promising direction. Despite its potential, the lack of large-scale multimodal data and the significant modality gap between video and radar pose substantial challenges in building robust videoradar RPM models. To handle these problems, we suggest leveraging unimodal pre-training and present the Spatial alignment and Temporal Matching (SATM) Adapter to effectively fine-tune pre-trained unimodal backbones into a multimodal RPM model. Given the distinct measurement principles of video-and radar-based methods, we propose Spatial Alignment to align the spatial distribution of their features. Furthermore, Temporal Matching is applied to mitigate waveform discrepancies between video and radar signals. By integrating these two modules into adapters, the unimodal backbones could retain their modality-specific knowledge while effectively extracting complementary features from each other. Extensive experiments across various challenging scenarios, including low light conditions and head motions, demonstrate that our approach significantly surpasses the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54552ebc-9f89-4b95-a789-086a3b1e154aBuilds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
- Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals MeasurementXin Liu, Josh Fromm, Shwetak N. Patel, Daniel McDuffNeurIPS 2020 · 436 citations
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference TransformerZitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao et al.CVPR 2022 · 255 citations
Related papers
- FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological MeasurementChenhang Ying, Huiyu Yang, Jieyi Ge, Zhaodong Sun et al.ICCV 2025 · 3 citations
- Blending camera and 77 GHz radar sensing for equitable, robust plethysmographyAlexander Vilesov, Pradyumna Chari, Adnan Armouti, Anirudh Bindiganavale Harish et al.SIGGRAPH 2022 · 29 citations
- Fusion-Vital: Video-RF Fusion Transformer for Advanced Remote Physiological MeasurementJae-Ho Choi, Ki-Bong Kang, Kyung-Tae KimAAAI 2024 · 20 citations
- Evidential Remote Physiological Measurement via Uncertainty-aware Fusion of Video and RFJieyi Ge, Zhaodong Sun, Wei Peng, Chenhang Ying et al.ACM MM 2025 · 1 citation
- Multimodal Emotion Recognition with Missing Modality via a Unified Multi-task Pre-training FrameworkZiyi Li, Wei-Long Zheng, Bao-Liang LuACM MM 2025 · 2 citations
