Video Semantic Segmentation via Sparse Temporal Transformer
Jiangtong Li, Wentao Wang, Junjie Chen, Li Niu, Jianlou Si, Chen Qian, Liqing Zhang
Abstract
Currently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bad34438-835e-4d57-bd45-66aee966a0cdCited by top-tier papers16
- Coarse-to-Fine Feature Mining for Video Semantic SegmentationGuolei Sun, Yun Liu, Henghui Ding, Thomas Probst et al.CVPR 2022 · 53 citations
- Mask Propagation for Efficient Video Semantic SegmentationYuetian Weng, Mingfei Han, Haoyu He, Mingjie Li et al.NeurIPS 2023 · 36 citations
- Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationYichen Yuan, Yifan Wang, Lijun Wang, Xiaoqi Zhao et al.ICCV 2023 · 16 citations
- PicT: A Slim Weakly Supervised Vision Transformer for Pavement Distress ClassificationWenhao Tang, Sheng Huang, Xiaoxian Zhang, Luwen HuangfuACM MM 2022 · 11 citations
- Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View BenchmarkWei Ji, Jingjing Li, Wenbo Li, Yilin Shen et al.NeurIPS 2024 · 8 citations
Related papers
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi et al.AAAI 2020 · 80 citations
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin et al.CVPR 2020
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi et al.CVPR 2021
- Temporally Efficient Vision Transformer for Video Instance SegmentationShusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang et al.CVPR 2022 · 68 citations
