Exploiting Temporal State Space Sharing for Video Semantic Segmentation
Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding, Jing Yang, Ender Konukoglu, Xue Geng, Xudong Jiang
Abstract
Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we introduce a Temporal Video State Space Sharing (TV3S) architecture to leverage Mamba state space models for temporal feature sharing. Our model features a selective gating mechanism that efficiently propagates relevant information across video frames, eliminating the need for a memory-heavy feature pool. By processing spatial patches independently and incorporating shifted operation, TV3S supports highly parallel computation in both training and inference stages, which reduces the delay in sequential state space processing and improves the scalability for long video sequences. Moreover, TV3S incorporates information from prior frames during inference, achieving long-range temporal coherence and superior adaptability to extended sequences. Evaluations on the VSPW and Cityscapes datasets reveal that our approach outperforms current state-of-the-art methods, establishing a new standard for VSS with consistent results across long video sequences. By achieving a good balance between accuracy and efficiency, TV3S shows a significant advancement in spatiotemporal modeling, paving the way for efficient video analysis. The code is publicly available at https://github.com/Ashesham/TV3S.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 583ffb9e-7b71-4cca-8e69-ccd149dde254Cited by top-tier papers3
- Generative Video Compression with One-Dimensional Latent RepresentationZihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li et al.CVPR 2026 · 5 citations
- Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time AdaptationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin YoonCVPR 2026 · 3 citations
- RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic SegmentationKai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan ZhouCVPR 2026 · 1 citation
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
Related papers
- VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal ModelingYamiao Ding, Tianrui Liu, Zhizhou Lu, Jun-Jie Huang et al.ACM MM 2025 · 1 citation
- Trajectory-aware Shifted State Space Models for Online Video Super-ResolutionQiang Zhu, Xiandong Meng, Yuxuan Jiang, Fan Zhang et al.ICLR 2026 · 3 citations
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentYachun Mi, Yu Li, Weicheng Meng, Chaofeng Chen et al.ICCV 2025 · 1 citation
- MaskViM: Domain Generalized Semantic Segmentation with State Space ModelsJiahao Li, Yang Lu, Yuan Xie, Yanyun QuAAAI 2025 · 1 citation
