CardioLive: Empowering Video Streaming with Online Cardiac Monitoring via Audio-Visual Learning
Sheng Lyu, Ruiming Huang, Sijie Ji, Yasar Abbas Ur Rehman, Lan Ma, Chenshu Wu
Abstract
Online Cardiac Monitoring (OCM) emerges as a compelling enhancement for the next-generation video streaming platforms. It enables various applications, including remote health, affective computing, and deepfake detection. Yet the physiological information encapsulated in the video streams has long been neglected. In this paper, we present the design and implementation of CardioLive, the first online cardiac monitoring system in video streaming platforms. We leverage the naturally co-existing video and audio streams and devise CardioNet, the first audio-visual network to learn the cardiac series. It incorporates multiple unique designs to extract temporal and spectral features, ensuring robust performance under realistic streaming conditions. To enable the Service-On-Demand OCM, we implement CardioLive as a plug-and-play middleware service and develop systematic solutions to practical issues including changing FPS and unsynchronized streams. Extensive evaluations demonstrate the effectiveness of our system. We achieve a Mean Squared Error of 1.79 BPM error, outperforming the videoonly and audio-only solutions by 69.2% and 81.2%, respectively. CardioLive achieves average throughput of 115.97 and 98.16 FPS in Zoom and YouTube. We believe our work opens up new applications for video stream systems. Code is available at https://github.com/aiot-lab/CardioLive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c947c8ad-8f3f-4da7-ac17-d5e9b8bf4fc1Builds on40
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals MeasurementXin Liu, Josh Fromm, Shwetak N. Patel, Daniel McDuffNeurIPS 2020 · 436 citations
- Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video EnhancementZitong Yu, Wei Peng, Xiaobai Li, Xiaopeng Hong et al.ICCV 2019 · 324 citations
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery et al.SIGCOMM 2020 · 238 citations
- DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsHua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie et al.ACM MM 2020 · 224 citations
Related papers
- Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera et al.ACM MM 2020 · 314 citations
- Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake VideosFazhong Liu, Yan Meng, Tian Dong, Guoxing Chen et al.CCS 2026
- LiveScreen: Video Chat Liveness Detection Leveraging Skin ReflectionHongbo Liu, Zhihua Li, Yucheng Xie, Ruizhe Jiang et al.INFOCOM 2020 · 13 citations
- AITransfer: Progressive AI-powered Transmission for Real-Time Point Cloud Video StreamingYakun Huang, Yuanwei Zhu, Xiuquan Qiao, Zhijie Tan et al.ACM MM 2021 · 34 citations
- EduLive: Re-Creating Cues for Instructor-Learners Interaction in Educational Live Streams with Learners' Transcript-Based AnnotationsJingchao Fang, Jeongeon Park, Juho Kim, Hao-Chuan WangCSCW 2024 · 1 citation
