Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the Edge
Changdi Yang, Pu Zhao, Yanyu Li, Wei Niu, Jiexiong Guan, Hao Tang, Minghai Qin, Bin Ren, Xue Lin, Yanzhi Wang
Abstract
With the ever-increasing popularity of edge devices, it is necessary to implement real-time segmentation on the edge for autonomous driving and many other applications. Vision Transformers (ViTs) have shown considerably stronger results for many vision tasks. However, ViTs with the fullattention mechanism usually consume a large number of computational resources, leading to difficulties for realtime inference on edge devices. In this paper, we aim to derive ViTs with fewer computations and fast inference speed to facilitate the dense prediction of semantic segmentation on edge devices. To achieve this, we propose a pruning parameterization method to formulate the pruning problem of semantic segmentation. Then we adopt a bi-level optimization method to solve this problem with the help of implicit gradients. Our experimental results demonstrate that we can achieve 38.9 mIoU on ADE20K val with a speed of 56.5 FPS on Samsung S21, which is the highest mIoU under the same computation constraint with real-time inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fa70e12-d6a2-4d8e-bf34-d4708b943b65Cited by top-tier papers10
- Search for Efficient Large Language ModelsXuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong et al.NeurIPS 2024 · 23 citations
- Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance AssessmentJun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang et al.AAAI 2025 · 23 citations
- Fast and Memory-Efficient Video Diffusion Using Streamlined InferenceZheng Zhan, Yushu Wu, Yifan Gong, Zichong Meng et al.NeurIPS 2024 · 23 citations
- Graph Lottery Ticket AutomatedGuibin Zhang, Kun Wang, Wei Huang, Yanwei Yue et al.ICLR 2024 · 17 citations
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness ConditionsLiuyuan Jiang, Quan Xiao, Lisha Chen, Tianyi ChenNeurIPS 2025 · 11 citations
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
Related papers
- Towards Real-Time Segmentation on the EdgeYanyu Li, Changdi Yang, Pu Zhao, Geng Yuan et al.AAAI 2023 · 19 citations
- Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesZhihao Shu, Xiaowei Yu, Zihao Wu, Wenqi Jia et al.NeurIPS 2024 · 6 citations
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie et al.HPCA 2023 · 117 citations
- TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationWenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen et al.CVPR 2022 · 313 citations
- Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-MemoryYi Li, Zijian Ye, Xiangqu Fu, Songqi Wang et al.DAC 2025 · 2 citations
