AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection
Long Chen, Hui Wang, Man Xu, Zexuan Li, Zizhu Fan
摘要
The You Only Look Once (YOLO) series has been a cornerstone in real-time object detection, renowned for its efficient convolutional design and rapid inference. However, its reliance on convolutional operations inherently limits its ability to capture long-range dependencies and rich contextual information, leading to suboptimal performance in complex scenes. Recently, State Space Models (SSM) have emerged as an efficient alternative to attention mechanisms, offering global representation with linear time complexity. In this paper, we propose AKCMamba-YOLO, a novel object detector that incorporates SSM into the YOLO architecture. We introduce 3CAKCMamba and 4CAKCMamba modules to a novel object detection framework, enabling enhanced channel interaction and cross-layer semantic fusion. This design improves multi-scale feature modeling while maintaining computational efficiency. To support safety-critical applications, we provide railway pedestrian detection dataset with 2,975 annotated images under complex scenarios. Experiments on COCO2017, Power Tower Foreign Object Detection Datasets, and our custom dataset show that AKCMamba-YOLO achieves superior accuracy and speed compared to state-of-the-art baselines, making it well-suited for real-time and resource-constrained environments. Code is available at https://github.com/ xlllchen/AKCMamba_YOLO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 被引用 845 次
相关 Paper
- Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelZeyu Wang, Chen Li, Huiying Xu, Xinzhong Zhu 等AAAI 2025 · 被引用 136 次
- Enabling True Global Perception in State Space Models for Visual TasksJie Hui, Zhenxiang Zhang, Wenyu Mi, Jianji WangICLR 2026
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object DetectionLanhu Wu, Zilin Gao, Hao Fei, Mong-Li Lee 等ACM MM 2025 · 被引用 3 次
- DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang 等CVPR 2026
- GroupMamba: Efficient Group-Based Visual State Space ModelAbdelrahman M. Shaker, Syed Talal Wasim, Salman H. Khan, Juergen Gall 等CVPR 2025
