Temporally Distributed Networks for Fast Video Semantic Segmentation
Ping Hu, Fabian Caba, Oliver Wang, Zhe Lin, Stan Sclaroff, Federico Perazzi
Abstract
We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower subnetworks. Leveraging the inherent temporal continuity in videos, we distribute these sub-networks over sequential frames. Therefore, at each time step, we only need to perform a lightweight computation to extract a sub-features group from a single sub-network. The full features used for segmentation are then recomposed by the application of a novel attention propagation module that compensates for geometry deformation between frames. A grouped knowledge distillation loss is also introduced to further improve the representation power at both full and sub-feature levels. Experiments on Cityscapes, CamVid, and NYUD-v2 demonstrate that our method achieves state-of-the-art accuracy with significantly faster speed and lower latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b522183a-197d-49ac-92ae-0f6fa7e1e288Cited by top-tier papers40
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai et al.NeurIPS 2021 · 92 citations
- Uncertainty-Aware Learning for Zero-Shot Semantic SegmentationPing Hu, Stan Sclaroff, Kate SaenkoNeurIPS 2020 · 79 citations
- Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationXiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen et al.CVPR 2022 · 71 citations
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li et al.CVPR 2022 · 58 citations
Builds on5
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- An Empirical Study of Spatial Attention Mechanisms in Deep NetworksXizhou Zhu, Dazhi Cheng, Zheng Zhang, Stephen Lin et al.ICCV 2019 · 522 citations
- Dynamic Multi-Scale Filters for Semantic SegmentationJunjun He, Zhongying Deng, Yu QiaoICCV 2019 · 287 citations
Related papers
- Ultrafast Video Attention Prediction with Coupled Knowledge DistillationKui Fu, Peipei Shi, Yafei Song, Shiming Ge et al.AAAI 2020 · 11 citations
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu et al.ACM MM 2021 · 47 citations
- Learning Structure Affinity for Video Depth EstimationYuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren et al.ACM MM 2021 · 12 citations
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click AnnotationsPeidong Liu, Zibin He, Xiyu Yan, Yong Jiang et al.ACM MM 2021 · 9 citations
- Fast Video Object Segmentation via Dynamic Targeting NetworkLu Zhang, Zhe Lin, Jianming Zhang, Huchuan Lu et al.ICCV 2019 · 59 citations
