Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, Jian Yang
摘要
One-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an individual prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the representations of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference (i.e., separately trained but compositely used in test) and (2) the inflexible Dirac delta distribution for localization when there is ambiguity and uncertainty which is often the case in complex scenes. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation of localization quality and classification, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain continuous labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the continuous version for successful optimization. On COCO test-dev, GFL achieves 45.0% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5%) and ATSS (43.6%) with higher or comparable inference speed, under the same backbone and training settings. Notably, our best model can achieve a single-model single-scale AP of 48.2%, at 10 FPS on a single 2080Ti GPU. Code and pretrained models are available at https://github.com/implus/GFocal .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper128
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen 等NeurIPS 2024 · 被引用 6,113 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 被引用 2,652 次
- TOOD: Task-aligned One-stage Object DetectionChengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott 等ICCV 2021 · 被引用 1,191 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
它引用的顶会 Paper10
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang 等ICCV 2019 · 被引用 1,056 次
- Scale-Aware Trident Networks for Object DetectionYanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 1,031 次
- Gaussian YOLOv3: An Accurate and Fast Object Detector Using Localization Uncertainty for Autonomous DrivingJiwoong Choi, Dayoung Chun, Hyun Kim, Hyuk-Jae LeeICCV 2019 · 被引用 445 次
相关 Paper
- DR Loss: Improving Object Detection by Distributional RankingQi Qian, Lei Chen, Hao Li, Rong JinCVPR 2020
- Reconcile Prediction Consistency for Balanced Object DetectionKeyang Wang, Lei ZhangICCV 2021 · 被引用 36 次
- VarifocalNet: An IoU-Aware Dense Object DetectorHaoyang Zhang, Ying Wang, Feras Dayoub, Niko SünderhaufCVPR 2021
- Ambiguity-Resistant Semi-Supervised Learning for Dense Object DetectionChang Liu, Weiming Zhang, Xiangru Lin, Wei Zhang 等CVPR 2023
- Enriched Feature Guided Refinement Network for Object DetectionJing Nie, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan 等ICCV 2019 · 被引用 80 次
