MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
Guanxiong Sun, Yang Hua, Guosheng Hu, Neil Robertson
Abstract
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) lightweight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both s peed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101. Code is available at https://github.com/guanxiongsun/vfe.pytorch .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 457f58e4-c96a-4fb4-8eca-8d863b7d992fCited by top-tier papers6
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- YOLOV: Making Still Image Object Detectors Great at Video Object DetectionYuheng Shi, Naiyan Wang, Xiaojie GuoAAAI 2023 · 83 citations
- Historical Test-time Prompt Tuning for Vision Foundation ModelsJingyi Zhang, Jiaxing Huang, Xiaoqin Zhang, Ling Shao et al.NeurIPS 2024 · 29 citations
- Spatio-temporal Prompting Network for Robust Video Feature ExtractionGuanxiong Sun, Chi Wang, Zhaoyu Zhang, Jiankang Deng et al.ICCV 2023 · 11 citations
- TGBFormer: Transformer-GraphFormer Blender Network for Video Object DetectionQiang Qi, Xiao WangAAAI 2025 · 5 citations
Builds on7
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- Progressive Sparse Local Attention for Video Object DetectionChaoxu Guo, Bin Fan, Jie Gu, Qian Zhang et al.ICCV 2019 · 95 citations
Related papers
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
- VIL-100: A New Dataset and A Baseline Model for Video Instance Lane DetectionYujun Zhang, Lei Zhu, Wei Feng, Huazhu Fu et al.ICCV 2021 · 67 citations
- Dual Semantic Fusion Network for Video Object DetectionLijian Lin, Haosheng Chen, Honglun Zhang, Jun Liang et al.ACM MM 2020 · 30 citations
- MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptYuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu et al.AAAI 2025 · 30 citations
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
