MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection
Mingjin Zhang, Yuanjun Ouyang, Fei Gao, Jie Guo, Qiming Zhang, Jing Zhang
Abstract
In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b8665c5-396b-4ce6-888f-8375acdf0c78Cited by top-tier papers1
Ask how each one uses itBuilds on14
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- ISNet: Shape Matters for Infrared Small Target DetectionMingjin Zhang, Rui Zhang, Yuxiang Yang, Haichen Bai et al.CVPR 2022 · 556 citations
Related papers
- CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared SequencesJingwen Ma, Xinpeng Zhang, Fan Shi, Xu Cheng et al.ICML 2026
- IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target DetectionMingjin Zhang, Xiaolong Li, Fei Gao, Jie GuoAAAI 2025 · 16 citations
- Explore Hybrid Modeling for Moving Infrared Small Target DetectionMingjin Zhang, Shilong Liu, Yuanjun Ouyang, Jie Guo et al.ACM MM 2024 · 6 citations
- Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target DetectionShengjia Chen, Luping Ji, Weiwei Duan, Shuang Peng et al.AAAI 2025 · 28 citations
- SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image PretrainingMingjin Zhang, Xiaolong Li, Fei Gao, Jie Guo et al.CVPR 2025
