Pixel-level Quality Assessment for Oriented Object Detection
Yunhui Zhu, Buliao Huang
Abstract
Modern oriented object detectors typically predict a set of bounding boxes and select the top-ranked ones based on estimated localization quality. Achieving high detection performance requires that the estimated quality closely aligns with the actual localization accuracy. To this end, existing approaches predict the Intersection over Union (IoU) between the predicted and ground-truth (GT) boxes as a proxy for localization quality. However, box-level IoU prediction suffers from a structural coupling issue: since the predicted box is derived from the detector’s internal estimation of the GT box, the predicted IoU—based on their similarity—can be overestimated for poorly localized boxes. To overcome this limitation, we propose a novel Pixel-level Quality Assessment (PQA) framework, which replaces box-level IoU prediction with the integration of pixel-level spatial consistency. PQA measures the alignment between each pixel’s relative position to the predicted box and its corresponding position to the GT box. By operating at the pixel level, PQA avoids directly comparing the predicted box with the estimated GT box, thereby eliminating the inherent similarity bias in box-level IoU prediction. Furthermore, we introduce a new integration metric that aggregates pixel-level spatial consistency into a unified quality score, yielding a more accurate approximation of the actual localization quality. Extensive experiments on HRSC2016 and DOTA demonstrate that PQA can be seamlessly integrated into various oriented object detectors, consistently improving performance (e.g., +5.96% AP50:95 on Rotated RetinaNet and +2.32% on STD).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a810365-3a84-4086-8676-72e5d94408e4Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- Oriented R-CNN for Object DetectionXingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao et al.ICCV 2021 · 1,070 citations
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler DivergenceXue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming et al.NeurIPS 2021 · 603 citations
Related papers
- Union-over-Intersections: Object Detection beyond Winner-Takes-AllAritra Bhowmik, Pascal Mettes, Martin R. Oswald, Cees G. M. SnoekICLR 2025
- Dynamic Anchor Learning for Arbitrary-Oriented Object DetectionQi Ming, Zhiqiang Zhou, Lingjuan Miao, Hongwei Zhang et al.AAAI 2021 · 332 citations
- Decoupled IoU Regression for Object DetectionYan Gao, Qimeng Wang, Xu Tang, Haochen Wang et al.ACM MM 2021 · 25 citations
- Polygon-to-Polygon Distance Loss for Rotated Object DetectionYang Yang, Jifeng Chen, Xiaopin Zhong, Yuanlong DengAAAI 2022 · 21 citations
- Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object DetectionXiang Li, Wenhai Wang, Xiaolin Hu, Jun Li et al.CVPR 2021
