YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors
Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao
Abstract
YOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 160 FPS and has the highest accuracy 56.8% AP among all known real-time object detectors with 30 FPS or higher on GPU V100. YOLOv7-E6 object detector (56 FPS V100, 55.9% AP) outperforms both transformer-based detector SWIN-L Cascade-Mask R-CNN (9.2 FPS A100, 53.9% AP) by 509% in speed and 2% in accuracy, and convolutionalbased detector ConvNeXt-XL Cascade-Mask R-CNN (8.6 FPS A100, 55.2% AP) by 551% in speed and 0.7% AP in accuracy, as well as YOLOv7 outperforms: YOLOR, YOLOX, Scaled-YOLOv4, YOLOv5, DETR, Deformable DETR, DINO-5scale-R50, ViT-Adapter-B and many other object detectors in speed and accuracy. Moreover, we train YOLOv7 only on MS COCO dataset from scratch without using any other datasets or pre-trained weights. Source code is released in https:// github.com/ WongKinYiu/ yolov7.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1dcd98-2b87-4931-a927-64d5e2207115Cited by top-tier papers12
- Exploring the Opportunities of AR for Enriching Storytelling with Family Photos between Grandparents and GrandchildrenZisu Li, Li Feng, Chen Liang, Yuru Huang et al.UbiComp 2023 · 35 citations
- Prompt Me Up: Unleashing the Power of Alignments for Multimodal Entity and Relation ExtractionXuming Hu, Junzhe Chen, Aiwei Liu, Shiao Meng et al.ACM MM 2023 · 30 citations
- MR Object Identification and Interaction: Fusing Object Situation Information from Heterogeneous SourcesJannis Strecker, Khakim Akhunov, Federico Carbone, Kimberly García et al.UbiComp 2023 · 15 citations
- R-TOSS: A Framework for Real-Time Object Detection using Semi-Structured PruningAbhishek Balasubramaniam, Febin Sunny, Sudeep PasrichaDAC 2023 · 14 citations
- BiPer: Binary Neural Networks Using a Periodic FunctionEdwin Vargas, Claudia V. Correa P., Carlos Hinojosa, Henry ArguelloCVPR 2024 · 10 citations
Builds on43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
Related papers
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- DynamicDet: A Unified Dynamic Architecture for Object DetectionZhihao Lin, Yongtao Wang, Jinhe Zhang, Xiaojie ChuCVPR 2023
- Scaled-YOLOv4: Scaling Cross Stage Partial NetworkChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2021
- YOLO-ULM: Ultra-Lightweight Models for Real-Time Object DetectionShasha Han, Chong Li, Xinning Wang, Xuebo LiCVPR 2026
- ViDT: An Efficient and Effective Fully Transformer-based Object DetectorHwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani et al.ICLR 2022 · 96 citations
