Efficient Semantic Segmentation by Altering Resolutions for Compressed Videos
Yubin Hu, Yuze He, Yanghao Li, Jisheng Li, Yuxing Han, Jiangtao Wen, Yong-Jin Liu
摘要
Video semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 backbone while maintaining high segmentation accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Compressed Video Prompt TuningBing Li, Jiaxin Chen, Xiuguo Bao, Di HuangNeurIPS 2023 · 被引用 11 次
- AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationHaifeng Zhong, Fan Tang, Zhuo Chen, Hyung Jin Chang 等ICCV 2025 · 被引用 9 次
- End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student LearningXin Yang, Wending Yan, Michael Bi Mi, Yuan Yuan 等NeurIPS 2024 · 被引用 6 次
- Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic SegmentationMatthieu Lin, Jenny Sheng, Yubin Hu, Yangguang Li 等AAAI 2024 · 被引用 3 次
- Efficient Motion-Aware Video MLLMZijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo 等CVPR 2025
它引用的顶会 Paper10
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Fast Object Detection in Compressed VideoShiyao Wang, Hongchao Lu, Zhidong DengICCV 2019 · 被引用 68 次
- Motion Adaptive Pose Estimation from Compressed VideosZhipeng Fan, Jun Liu, Yao WangICCV 2021 · 被引用 24 次
- BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online PoliciesThomas Verelst, Tinne TuytelaarsICCV 2021 · 被引用 18 次
相关 Paper
- Mask Propagation for Efficient Video Semantic SegmentationYuetian Weng, Mingfei Han, Haoyu He, Mingjie Li 等NeurIPS 2023 · 被引用 36 次
- Vanishing-Point-Guided Video Semantic Segmentation of Driving ScenesDiandian Guo, Deng-Ping Fan, Tongyu Lu, Christos Sakaridis 等CVPR 2024
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin 等CVPR 2020
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu 等ACM MM 2021 · 被引用 47 次
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo 等NeurIPS 2022 · 被引用 17 次
