Dual-Temporal Exemplar Representation Network for Video Semantic Segmentation
Xiaolong Xu, Lei Zhang, Jiayi Li, Lituan Wang, Yifan Guan, Yu Yan, Leyi Zhang, Hao Song
摘要
Video semantic segmentation aims to assign a class label for each pixel in every video frame. Existing methods predominantly follow the reference-target interaction paradigm, focusing on extracting local temporal contexts while neglecting the integration of global temporal information. Moreover, complex dynamics and varying lighting conditions introduce inter-frame intra-class discrepancies in feature representations, leading to unstable predictions. In this paper, we propose a novel framework, the Dual-Temporal Exemplar Representation Network (DTERN), which utilizes the strong representational capability of cluster centers, i.e., exemplars, to effectively model both local and global temporal information. DTERN consists of two core modules: 1) the Local Temporal Exemplar Module (LTEM), which constructs local exemplars to capture local temporal contexts, ensuring stable and reliable predictions. 2) the Global Temporal Exemplar Module (GTEM), which introduces learnable global exemplars to dynamically model global temporal information, thereby improving the effective consistency of segmentation. Furthermore, we observe that the existing Video Consistency (VC) metric fails to evaluate segmentation accuracy and lacks sensitivity to small-object segmentation. To this end, we propose Video Effective Consistency (VEC) to comprehensively evaluate temporal consistency and segmentation effectiveness. Experiments on VSPW and Cityscape demonstrate that DTERN outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang 等ICCV 2019 · 被引用 694 次
- ACFNet: Attentional Class Feature Network for Semantic SegmentationFan Zhang, Yanqin Chen, Zhihang Li, Zhibin Hong 等ICCV 2019 · 被引用 297 次
- Coarse-to-Fine Feature Mining for Video Semantic SegmentationGuolei Sun, Yun Liu, Henghui Ding, Thomas Probst 等CVPR 2022 · 被引用 53 次
相关 Paper
- Domain Adaptive Video Segmentation via Temporal Consistency RegularizationDayan Guan, Jiaxing Huang, Aoran Xiao, Shijian LuICCV 2021 · 被引用 44 次
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi 等AAAI 2020 · 被引用 80 次
- Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationJiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang 等CVPR 2023
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu 等ACM MM 2021 · 被引用 47 次
- Video-Text Pre-training with Learned Regions for RetrievalRui Yan, Mike Zheng Shou, Yixiao Ge, Jinpeng Wang 等AAAI 2023 · 被引用 29 次
