Real-Time, Accurate, and Consistent Video Semantic Segmentation via Unsupervised Adaptation and Cross-Unit Deployment on Mobile Device
Hyojin Park, Alan Yessenbayev, Tushar Singhal, Navin Kumar Adhikari, Yizhe Zhang, Shubhankar Mangesh Borse, Hong Cai, Frank Mayer, Balaji Calidas, Nilesh Prasad Pandey, Fei Yin, Fatih Porikli
Abstract
This demonstration showcases our innovations on efficient, accurate, and temporally consistent video semantic segmentation on mobile device. We employ our test-time unsupervised scheme, AuxAdapt, to enable the segmentation model to adapt to a given video in an online manner. More specifically, we leverage a small auxiliary network to perform weight updates and keep the large, main segmen-tation network frozen. This significantly reduces the computational cost of adaptation when compared to previous methods (e.g., Tent, DVP), and at the same time, prevents catastrophic forgetting. By running AuxAdapt, we can considerably improve the temporal consistency of video segmentation while maintaining the accuracy. We demonstrate how to efficiently deploy our adaptive video segmentation algorithm on a smartphone powered by a Snapdragon® Mobile Platform <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Snapdragon is a product of Qualcomm Technologies, Inc. and/or its subsidiaries., Rather than simply running the entire algorithm on the GPU, we adopt a crossunit deployment strategy. The main network, which will be frozen during test time, will perform inferences on a highly optimized AI accelerator unit, while the small auxiliary net-work, which will be updated on the fly, will run forward passes and back-propagations on the GPU. Such a deployment scheme best utilizes the available processing power on the smartphone and enables real-time operation of our adaptive video segmentation algorithm. We provide example videos in supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9d2870c-41d8-4857-a42c-5fb26e6328c5Cited by top-tier papers3
- SRoUDA: Meta Self-Training for Robust Unsupervised Domain AdaptationWanqing Zhu, Jia-Li Yin, Bo-Hao Chen, Ximeng LiuAAAI 2023 · 14 citations
- Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time AdaptationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin YoonCVPR 2026 · 3 citations
- DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive SegmentationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Wooseong Jeong et al.ICCV 2025 · 1 citation
Builds on2
Related papers
- Real-Time Video Inference on Edge Devices via Adaptive Model StreamingMehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad AlizadehICCV 2021 · 58 citations
- SwiftNet: Real-Time Video Object SegmentationHaochen Wang, Xiaolong Jiang, Haibing Ren, Yao Hu et al.CVPR 2021
- MobileInst: Video Instance Segmentation on the MobileRenhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang et al.AAAI 2024 · 10 citations
- VR-DANN: Real-Time Video Recognition via Decoder-Assisted Neural Network AccelerationZhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke et al.MICRO 2020 · 29 citations
- Domain Adaptive Video Segmentation via Temporal Consistency RegularizationDayan Guan, Jiaxing Huang, Aoran Xiao, Shijian LuICCV 2021 · 44 citations
