Every Frame Counts: Joint Learning of Video Segmentation and Optical Flow
Mingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi, Zhiwu Lu, Ping Luo
Abstract
A major challenge for video semantic segmentation is the lack of labeled data. In most benchmark datasets, only one frame of a video clip is annotated, which makes most supervised methods fail to utilize information from the rest of the frames. To exploit the spatio-temporal information in videos, many previous works use pre-computed optical flows, which encode the temporal consistency to improve the video segmentation. However, the video segmentation and optical flow estimation are still considered as two separate tasks. In this paper, we propose a novel framework for joint video semantic segmentation and optical flow estimation. Semantic segmentation brings semantic information to handle occlusion for more robust optical flow estimation, while the non-occluded optical flow provides accurate pixel-level temporal correspondences to guarantee the temporal consistency of the segmentation. Moreover, our framework is able to utilize both labeled and unlabeled frames in the video through joint training, while no additional calculation is required in inference. Extensive experiments show that the proposed model makes the video semantic segmentation and optical flow estimation benefit from each other and outperforms existing methods under the same settings in both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2d13bd4-ac24-459d-b632-dc1b19990887Cited by top-tier papers18
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- Domain Adaptive Video Segmentation via Temporal Consistency RegularizationDayan Guan, Jiaxing Huang, Aoran Xiao, Shijian LuICCV 2021 · 44 citations
- Learning Versatile Neural Architectures by Propagating Network CodesMingyu Ding, Yuqi Huo, Haoyu Lu, Linjie Yang et al.ICLR 2022 · 14 citations
- Semi-Supervised Video Semantic Segmentation with Inter-Frame Feature ReconstructionJiafan Zhuang, Zilei Wang, Yuan GaoCVPR 2022 · 14 citations
- SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous DrivingShuai Yuan, Shuzhi Yu, Hannah Kim, Carlo TomasiICCV 2023 · 13 citations
Related papers
- Unsupervised Space-Time Network for Temporally-Consistent Segmentation of Multiple MotionsEtienne Meunier, Patrick BouthemyCVPR 2023
- Exploit Domain-Robust Optical Flow in Domain Adaptive Video Semantic SegmentationYuan Gao, Zilei Wang, Jiafan Zhuang, Yixin Zhang et al.AAAI 2023 · 12 citations
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu et al.ACM MM 2021 · 47 citations
- OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware InterpolationJisoo Jeong, Hong Cai, Risheek Garrepalli, Jamie Menjay Lin et al.CVPR 2024
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong et al.CVPR 2022 · 83 citations
