Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
Haipeng Zhou, Hongqiu Wang, Tian Ye, Zhaohu Xing, Jun Ma, Ping Li, Qiong Wang, Lei Zhu
Abstract
Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at https://github.com/haipengzhou856/TBGDiff.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5af3e4d-649b-4b10-8380-7db0a6a75aedCited by top-tier papers5
- Language-Driven Interactive Shadow DetectionHongqiu Wang, Wei Wang, Haipeng Zhou, Huihui Xu et al.ACM MM 2024 · 7 citations
- Dynamic Shadow Unveils Invisible Semantics for Video OutpaintingRuilin Li, Hang Yu, Jiayan QiuNeurIPS 2025 · 2 citations
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingZhicheng Li, Kunyang Sun, Rui Yao, Hancheng Zhu et al.AAAI 2026
- SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationJianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye et al.CVPR 2025
- Toward Real-World High-Precision Image Matting and SegmentationHaipeng Zhou, Zhaohu Xing, Hongqiu Wang, Jun Ma et al.AAAI 2026
Builds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- SCOTCH and SODA: A Transformer Video Shadow Detection FrameworkLihao Liu, Jean Prost, Lei Zhu, Nicolas Papadakis et al.CVPR 2023
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- Video Shadow Detection via Spatio-Temporal Interpolation Consistency TrainingXiao Lu, Yihong Cao, Sheng Liu, Chengjiang Long et al.CVPR 2022 · 24 citations
- SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow DetectionRunmin Cong, Yuchen Guan, Jinpeng Chen, Wei Zhang et al.ACM MM 2023 · 15 citations
- Triple-Cooperative Video Shadow DetectionZhihao Chen, Liang Wan, Lei Zhu, Jia Shen et al.CVPR 2021
