Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training
Xiao Lu, Yihong Cao, Sheng Liu, Chengjiang Long, Zipei Chen, Xuanyu Zhou, Yimin Yang, Chunxia Xiao
Abstract
It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Spatio-Temporal Interpolation Consistency Training (STICT) framework to rationally feed the unlabeled video frames together with the labeled images into an image shadow detection network training. Specifically, we propose the Spatial and Temporal ICT, in which we define two new interpolation schemes, i.e., the spatial interpolation and the temporal interpolation. We then derive the spatial and temporal interpolation consistency constraints accordingly for enhancing generalization in the pixel-wise classification task and for encouraging temporal consistent predictions, respectively. In addition, we design a Scale-Aware Network for multi-scale shadow knowledge learning in images, and propose a scale-consistency constraint to minimize the discrepancy among the predictions at different scales. Our proposed approach is extensively validated on the ViSha dataset and a self-annotated dataset. Experimental results show that, even without video labels, our approach is better than most state of the art supervised, semi-supervised or unsupervised image/video shadow detection methods and other methods in related tasks. Code and dataset are available at https://github.com/ yihong-97/STICT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ffeab63-ec17-4783-928a-494e83167762Cited by top-tier papers7
- Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionHaipeng Zhou, Hongqiu Wang, Tian Ye, Zhaohu Xing et al.ACM MM 2024 · 18 citations
- SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow DetectionRunmin Cong, Yuchen Guan, Jinpeng Chen, Wei Zhang et al.ACM MM 2023 · 15 citations
- Multi-view Spectral Polarization Propagation for Video Glass SegmentationYu Qiao, Bo Dong, Ao Jin, Yu Fu et al.ICCV 2023 · 9 citations
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingZhicheng Li, Kunyang Sun, Rui Yao, Hancheng Zhu et al.AAAI 2026
- Learning to Detect Mirrors from Videos via Dual CorrespondencesJiaying Lin, Xin Tan, Rynson W. H. LauCVPR 2023
Builds on13
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and RemovalBin Ding, Chengjiang Long, Ling Zhang, Chunxia XiaoICCV 2019 · 171 citations
- Semi-Supervised Video Salient Object Detection Using Pseudo-LabelsPengxiang Yan, Guanbin Li, Yuan Xie, Zhen Li et al.ICCV 2019 · 134 citations
- CANet: A Context-Aware Network for Shadow RemovalZipei Chen, Chengjiang Long, Ling Zhang, Chunxia XiaoICCV 2021 · 118 citations
- RIS-GAN: Explore Residual and Illumination with Generative Adversarial Networks for Shadow RemovalLing Zhang, Chengjiang Long, Xiaolong Zhang, Chunxia XiaoAAAI 2020 · 106 citations
Related papers
- Semi-supervised Video Shadow Detection via Image-assisted Pseudo-label GenerationZipei Chen, Xiao Lu, Ling Zhang, Chunxia XiaoACM MM 2022 · 9 citations
- Triple-Cooperative Video Shadow DetectionZhihao Chen, Liang Wan, Lei Zhu, Jia Shen et al.CVPR 2021
- A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionZhihao Chen, Lei Zhu, Liang Wan, Song Wang et al.CVPR 2020
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 31 citations
- Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark TrackingCongcong Zhu, Xiaoqiang Li, Jide Li, Guangtai Ding et al.ACM MM 2020 · 10 citations
