SCOTCH and SODA: A Transformer Video Shadow Detection Framework
Lihao Liu, Jean Prost, Lei Zhu, Nicolas Papadakis, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
Abstract
Shadows in videos are difficult to detect because of the large shadow deformation between frames. In this work, we argue that accounting for shadow deformation is essential when designing a video shadow detection method. To this end, we introduce the shadow deformation attention trajectory (SODA), a new type of video self-attention module, specially designed to handle the large shadow deformations in videos. Moreover, we present a new shadow contrastive learning mechanism (SCOTCH) which aims at guiding the network to learn a unified shadow representation from massive positive shadow pairs across different videos. We demonstrate empirically the effectiveness of our two contributions in an ablation study. Furthermore, we show that SCOTCH and SODA significantly outperforms existing techniques for video shadow detection. Code is
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc3f1f24-8d43-481c-9f04-9a07ff721ec9Cited by top-tier papers9
- Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionHaipeng Zhou, Hongqiu Wang, Tian Ye, Zhaohu Xing et al.ACM MM 2024 · 18 citations
- Language-Driven Interactive Shadow DetectionHongqiu Wang, Wei Wang, Haipeng Zhou, Huihui Xu et al.ACM MM 2024 · 7 citations
- Controllable-Lpmoe: Adapting to Challenging Object Segmentation Via Dynamic Local Priors From Mixture-Of-ExpertsYanguang Sun, Jiawei Lian, Jian Yang, Lei LuoICCV 2025 · 4 citations
- Dynamic Shadow Unveils Invisible Semantics for Video OutpaintingRuilin Li, Hang Yu, Jiayan QiuNeurIPS 2025 · 2 citations
- TrafficMOT: A Challenging Dataset for Multi-Object Tracking in Complex Traffic ScenariosLihao Liu, Yanqi Cheng, Zhongying Deng, Shujun Wang et al.ACM MM 2024 · 2 citations
Builds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
Related papers
- Triple-Cooperative Video Shadow DetectionZhihao Chen, Liang Wan, Lei Zhu, Jia Shen et al.CVPR 2021
- Video Shadow Detection via Spatio-Temporal Interpolation Consistency TrainingXiao Lu, Yihong Cao, Sheng Liu, Chengjiang Long et al.CVPR 2022 · 24 citations
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingZhicheng Li, Kunyang Sun, Rui Yao, Hancheng Zhu et al.AAAI 2026
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong et al.CVPR 2022 · 83 citations
- Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationSeokju Lee, François Rameau, Fei Pan, In So KweonICCV 2021 · 38 citations
