Towards Real-time Video Compressive Sensing on Mobile Devices
Miao Cao, Lishun Wang, Huan Wang, Guoqing Wang, Xin Yuan
Abstract
Video Snapshot Compressive Imaging (SCI) uses a low-speed 2D camera to capture high-speed scenes as snapshot compressed measurements, followed by a reconstruction algorithm to retrieve the high-speed video frames. The fast evolving mobile devices and existing high-performance video SCI reconstruction algorithms motivate us to develop mobile reconstruction methods for real-world applications. Yet, it is still challenging to deploy previous reconstruction algorithms on mobile devices due to the complex inference process, let alone real-time mobile reconstruction. To the best of our knowledge, there is no video SCI reconstruction model designed to run on the mobile devices. Towards this end, in this paper, we present an effective approach for video SCI reconstruction, dubbed MobileSCI, which can run at real-time speed on the mobile devices for the first time. Specifically, we first build a U-shaped 2D convolution-based architecture, which is much more efficient and mobile-friendly than previous state-of-the-art reconstruction methods. Besides, an efficient feature mixing block, based on the channel splitting and shuffling mechanisms, is introduced as a novel bottleneck block of our proposed MobileSCI to alleviate the computational burden. Finally, a customized knowledge distillation strategy is utilized to further improve the reconstruction quality. Extensive results on both simulated and real data show that our proposed MobileSCI can achieve superior reconstruction quality with high efficiency on the mobile devices. Particularly, we can reconstruct a 256x256x8 snapshot compressed measurement with real-time performance (about 35 FPS) on an iPhone 15. Code is available at https://github.com/mcao92/MobileSCI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61745452-e284-4272-8257-0f8daf8129fcCited by top-tier papers1
Ask how each one uses itBuilds on15
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- GhostNetV2: Enhance Cheap Operation with Long-Range AttentionYehui Tang, Kai Han, Jianyuan Guo, Chang Xu et al.NeurIPS 2022 · 634 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis et al.ICCV 2023 · 300 citations
Related papers
- EfficientSCI: Densely Connected Network with Space-time Factorization for Large-scale Video Snapshot Compressive ImagingLishun Wang, Miao Cao, Xin YuanCVPR 2023
- Memory-Efficient Network for Large-Scale Video Compressive SensingZiheng Cheng, Bo Chen, Guanliang Liu, Hao Zhang et al.CVPR 2021
- MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive SensingZhengjue Wang, Hao Zhang, Ziheng Cheng, Bo Chen et al.CVPR 2021
- DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive ImagingXingjian Jiang, Lishun Wang, Ping Wang, Xin YuanCVPR 2026
- MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive ImagingZhenghao Pan, Haijin Zeng, Jiezhang Cao, Yongyong Chen et al.NeurIPS 2024 · 12 citations
