LSTFE-Net: Long Short-Term Feature Enhancement Network for Video Small Object Detection
Jinsheng Xiao, Yuanxu Wu, Yunhua Chen, Shurui Wang, Zhongyuan Wang, Jiayi Ma
Abstract
Video small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropriate features. Since information from frames at different positions contributes differently to small objects, it is not ideal to assume that using one universal method will extract proper features. We find that context information from the long-term frame and temporal information from the short-term frame are two useful cues for video small object detection. To fully utilize these two cues, we propose a long short-term feature enhancement network (LSTFE-Net) for video small object detection. First, we develop a plugand-play spatio-temporal feature alignment module to create temporal correspondences between the short-term and current frames. Then, we propose a frame selection module to select the long-term frame that can provide the most additional context information. Finally, we propose a long short-term feature aggregation module to fuse long shortterm features. Compared to other state-of-the-art methods, our LSTFE-Net achieves 4.4% absolute boosts in AP on the FL-Drones dataset. More details can be found at https://github.com/xiaojs18/LSTFE-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fbf9f6b-2a18-4666-9c45-0f32840559f5Cited by top-tier papers4
- MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target DetectionMingjin Zhang, Yuanjun Ouyang, Fei Gao, Jie Guo et al.AAAI 2025 · 10 citations
- ESOD: Event-Based Small Object DetectionQuanmin Liang, Jinyi Lu, Qiang Li, Shuai Liu et al.ACM MM 2025 · 2 citations
- Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target DetectionHouzhang Fang, Shukai Guo, Qiuhuan Chen, Yi Chang et al.AAAI 2026
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
Builds on9
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object DetectionChenhongyi Yang, Zehao Huang, Naiyan WangCVPR 2022 · 472 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 171 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
Related papers
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- Alleviating Spatial Misalignment and Motion Interference for UAV-based Video RecognitionGege Shi, Xueyang Fu, Chengzhi Cao, Zheng-Jun ZhaACM MM 2023 · 8 citations
- Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationJiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang et al.CVPR 2023
- Dogfight: Detecting Drones From Drones VideosMuhammad Waseem Ashraf, Waqas Sultani, Mubarak ShahCVPR 2021
