YOLOv12: Attention-Centric Real-Time Object Detectors
Yunjie Tian, Qixiang Ye, David S. Doermann
摘要
Enhancing the network architecture of the YOLO framework has been crucial for a long time, but has focused on CNN-based improvements despite the proven superiority of attention mechanisms in modeling capabilities. This is because attention-based models cannot match the speed of CNN-based models. This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. YOLOv12 surpasses all popular real-time object detectors in accuracy with competitive speed. For example, YOLOv12-N achieves 40.6% mAP with an inference latency of 1.64 ms on a T4 GPU, outperforming advanced YOLOv10-N / YOLOv11-N by 2.1%/1.2% mAP with a comparable speed. This advantage extends to other model scales. YOLOv12 also surpasses end-to-end real-time detectors that improve DETR, such as RT-DETR / RT-DETRv2: YOLOv12-S beats RT-DETR-R18 / RT-DETRv2-R18 while running 42% faster, using only 36% of the computation and 45% of the parameters. More comparisons are shown in Figure 1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object DetectionZhangchi Hu, Peixi Wu, Jie Chen, Huyue Zhu 等ACM MM 2025 · 被引用 25 次
- YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time DetectionXu Lin, Jinlong Peng, Zhenye Gan, Jiawen Zhu 等CVPR 2026 · 被引用 22 次
- 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree ViewsXiaobiao Du, Yida Wang, Haiyang Sun, Zhuojie Wu 等ICCV 2025 · 被引用 10 次
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationYuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa 等CVPR 2026 · 被引用 10 次
- MoCha: End-to-End Video Character Replacement without Structural GuidanceZhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng 等CVPR 2026 · 被引用 9 次
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- YOLO-ULM: Ultra-Lightweight Models for Real-Time Object DetectionShasha Han, Chong Li, Xinning Wang, Xuebo LiCVPR 2026
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen 等NeurIPS 2024 · 被引用 6,113 次
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object DetectorsChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2023
- Scaled-YOLOv4: Scaling Cross Stage Partial NetworkChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2021
