Towards Real-Time Segmentation on the Edge
Yanyu Li, Changdi Yang, Pu Zhao, Geng Yuan, Wei Niu, Jiexiong Guan, Hao Tang, Minghai Qin, Qing Jin, Bin Ren, Xue Lin, Yanzhi Wang
Abstract
The research in real-time segmentation mainly focuses on desktop GPUs. However, autonomous driving and many other applications rely on real-time segmentation on the edge, and current arts are far from the goal. In addition, recent advances in vision transformers also inspire us to re-design the network architecture for dense prediction task. In this work, we propose to combine the self attention block with lightweight convolutions to form new building blocks, and employ latency constraints to search an efficient sub-network. We train an MLP latency model based on generated architecture configurations and their latency measured on mobile devices, so that we can predict the latency of subnets during search phase. To the best of our knowledge, we are the first to achieve over 74% mIoU on Cityscapes with semi-real-time inference (over 15 FPS) on mobile GPU from an off-the-shelf phone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1911f84e-6aa0-4c28-a3ba-a225772d76b6Cited by top-tier papers5
- Search for Efficient Large Language ModelsXuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong et al.NeurIPS 2024 · 23 citations
- Sparse Learning for State Space Models on MobileXuan Shen, Hangyu Zheng, Yifan Gong, Zhenglun Kong et al.ICLR 2025
- Searching Efficient Semantic Segmentation Architectures via Dynamic Path SelectionYuxi Liu, Min Liu, Shuai Jiang, Yi Tang et al.NeurIPS 2025
- Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the EdgeChangdi Yang, Pu Zhao, Yanyu Li, Wei Niu et al.CVPR 2023
- QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgeXuan Shen, Weize Ma, Jing Liu, Changdi Yang et al.CVPR 2025
Builds on12
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 362 citations
Related papers
- Graph-Guided Architecture Search for Real-Time Semantic SegmentationPeiwen Lin, Peng Sun, Guangliang Cheng, Sirui Xie et al.CVPR 2020
- MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile DevicesBo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin et al.CVPR 2020
- MobileDets: Searching for Object Detection Architectures for Mobile AcceleratorsYunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin et al.CVPR 2021
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue et al.NeurIPS 2022 · 30 citations
- TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationWenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen et al.CVPR 2022 · 313 citations
