Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target Detection
Houzhang Fang, Shukai Guo, Qiuhuan Chen, Yi Chang, Luxin Yan
摘要
Moving infrared small target detection (IRSTD) plays a critical role in practical applications, such as surveillance of unmanned aerial vehicles (UAVs) and UAV-based search system. Moving IRSTD still remains highly challenging due to weak target features and complex background interference. Accurate spatio-temporal feature modeling is crucial for moving target detection, typically achieved through either temporal differences or spatio-temporal (3D) convolutions. Temporal difference can explicitly leverage motion cues but exhibits limited capability in extracting spatial features, whereas 3D convolution effectively represents spatio-temporal features yet lacks explicit awareness of motion dynamics along the temporal dimension. In this paper, we propose a novel moving IRSTD network (TDCNet), which effectively extracts and enhances spatio-temporal features for accurate target detection. Specifically, we introduce a novel temporal difference convolution (TDC) re-parameterization module that comprises three parallel TDC blocks designed to capture contextual dependencies across different temporal ranges. Each TDC block fuses temporal difference and 3D convolution into a unified spatio-temporal convolution representation. This re-parameterized module can effectively capture multi-scale motion contextual features while suppressing pseudo-motion clutter in complex backgrounds, significantly improving detection performance. Moreover, we propose a TDC-guided spatio-temporal attention mechanism that performs cross-attention between the spatio-temporal features extracted from the TDC-based backbone and a parallel 3D backbone. This mechanism models their global semantic dependencies to refine the current frame’s features, thereby guiding the model to focus more accurately on critical target regions. To facilitate comprehensive evaluation, we construct a new challenging benchmark, IRSTD-UAV, consisting of 15,106 real infrared images with diverse low signal-to-clutter ratio scenarios and complex backgrounds. Extensive experiments on IRSTD-UAV and public infrared datasets demonstrate that our TDCNet achieves state-of-the-art detection performance in moving target detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Pinwheel-shaped Convolution and Scale-based Dynamic Loss for Infrared Small Target DetectionJiangnan Yang, Shuangli Liu, Jingjun Wu, Xinyu Su 等AAAI 2025 · 被引用 176 次
- IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel PruningMingjin Zhang, Handi Yang, Jie Guo, Yunsong Li 等AAAI 2024 · 被引用 159 次
相关 Paper
- ISNet: Shape Matters for Infrared Small Target DetectionMingjin Zhang, Rui Zhang, Yuxiang Yang, Haichen Bai 等CVPR 2022 · 被引用 556 次
- MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target DetectionMingjin Zhang, Yuanjun Ouyang, Fei Gao, Jie Guo 等AAAI 2025 · 被引用 10 次
- Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target DetectionHouzhang Fang, Xiaolin Wang, Zengyang Li, Lu Wang 等CVPR 2025
- Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex ScenesFeng Huang, Shuyuan Zheng, Zhaobing Qiu, Huanxian Liu 等ICCV 2025 · 被引用 3 次
- Explore Hybrid Modeling for Moving Infrared Small Target DetectionMingjin Zhang, Shilong Liu, Yuanjun Ouyang, Jie Guo 等ACM MM 2024 · 被引用 6 次
