MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection
Mingjin Zhang, Yuanjun Ouyang, Fei Gao, Jie Guo, Qiming Zhang, Jing Zhang
摘要
In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu 等NeurIPS 2021 · 被引用 798 次
- ISNet: Shape Matters for Infrared Small Target DetectionMingjin Zhang, Rui Zhang, Yuxiang Yang, Haichen Bai 等CVPR 2022 · 被引用 556 次
相关 Paper
- CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared SequencesJingwen Ma, Xinpeng Zhang, Fan Shi, Xu Cheng 等ICML 2026
- IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target DetectionMingjin Zhang, Xiaolong Li, Fei Gao, Jie GuoAAAI 2025 · 被引用 16 次
- Explore Hybrid Modeling for Moving Infrared Small Target DetectionMingjin Zhang, Shilong Liu, Yuanjun Ouyang, Jie Guo 等ACM MM 2024 · 被引用 6 次
- Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target DetectionShengjia Chen, Luping Ji, Weiwei Duan, Shuang Peng 等AAAI 2025 · 被引用 28 次
- SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image PretrainingMingjin Zhang, Xiaolong Li, Fei Gao, Jie Guo 等CVPR 2025
