YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection
Shasha Han, Chong Li, Xinning Wang, Xuebo Li
Abstract
YOLO series lead object detection with superior accuracy and speed. However, both convolutional and self-attention based architectures suffer from parameter redundancy and insufficient computational efficiency. Existing lightweight methods excessively pursue speed while ignoring the loss of important information during feature extraction and spatial transformation across different stages. Thus, effective lightweighting is crucial for detection performance. We propose YOLO-ULM, an ultra-lightweight real-time detector that achieves accelerated inference while preserving high accuracy. We innovatively design a variety of dual efficiency- and accuracy-driven modules, including efficient feature aggregation and multi-scale downsampling modules, as well as a more focused complete-IoU loss function. To validate our approach, we train it from scratch on COCO dataset without pretrained weights. By refining backbone parameters, we extend it to YOLO-ULM-Turbo for accelerated inference. YOLO-ULM surpasses state-of-the-art real-time detectors like YOLOv11/YOLOv12/YOLOv13 and RT-DETR. On a T4 GPU, YOLO-ULM-N achieves 41.6% mAP with an inference latency of 1.52 ms, outperforming YOLOv11-N (2.2%) and YOLOv12-N (1.0%). YOLO-ULM-S exceeds RT-DETR-R18 by 1.6% mAP with 64.7% fewer FLOPs and 63% fewer parameters. YOLO-ULM-L / X surpass YOLOv13-L / X by 0.7% and 0.8% respectively in mAP. YOLO-ULM-Turbo matches YOLOv12-Turbo's performance but uses less computation, with Turbo-N variant achieving 0.3% higher mAP and 16% fewer parameters than YOLOv12-Turbo-N.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce701e6c-0db3-4e93-baa4-9ff41feaeeb4Builds on7
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li et al.AAAI 2020 · 4,823 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
Related papers
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 2,652 citations
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen et al.NeurIPS 2024 · 6,113 citations
- DEIM: DETR with Improved Matching for Fast ConvergenceShihua Huang, Zhichao Lu, Xiaodong Cun, Yongjun Yu et al.CVPR 2025
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object DetectorsChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2023
- Gold-YOLO: Efficient Object Detector via Gather-and-Distribute MechanismChengcheng Wang, Wei He, Ying Nie, Jianyuan Guo et al.NeurIPS 2023 · 732 citations
