MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
Guanxiong Sun, Yang Hua, Guosheng Hu, Neil Robertson
摘要
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) lightweight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both s peed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101. Code is available at https://github.com/guanxiongsun/vfe.pytorch .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu 等ACM MM 2021 · 被引用 106 次
- YOLOV: Making Still Image Object Detectors Great at Video Object DetectionYuheng Shi, Naiyan Wang, Xiaojie GuoAAAI 2023 · 被引用 83 次
- Historical Test-time Prompt Tuning for Vision Foundation ModelsJingyi Zhang, Jiaxing Huang, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 29 次
- Spatio-temporal Prompting Network for Robust Video Feature ExtractionGuanxiong Sun, Chi Wang, Zhaoyu Zhang, Jiankang Deng 等ICCV 2023 · 被引用 11 次
- TGBFormer: Transformer-GraphFormer Blender Network for Video Object DetectionQiang Qi, Xiao WangAAAI 2025 · 被引用 5 次
它引用的顶会 Paper7
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 236 次
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou 等ICCV 2019 · 被引用 211 次
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang 等ICCV 2019 · 被引用 109 次
- Progressive Sparse Local Attention for Video Object DetectionChaoxu Guo, Bin Fan, Jie Gu, Qian Zhang 等ICCV 2019 · 被引用 95 次
相关 Paper
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
- VIL-100: A New Dataset and A Baseline Model for Video Instance Lane DetectionYujun Zhang, Lei Zhu, Wei Feng, Huazhu Fu 等ICCV 2021 · 被引用 67 次
- Dual Semantic Fusion Network for Video Object DetectionLijian Lin, Haosheng Chen, Honglun Zhang, Jun Liang 等ACM MM 2020 · 被引用 30 次
- MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptYuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu 等AAAI 2025 · 被引用 30 次
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
